Is Artificial Intelligence Accurate at Reading Blood Work?

What the published research actually shows, and how we built our report around it.

Anyone paying a machine to read their bloodwork deserves the honest answer before the sales page. Published research says that artificial intelligence, meaning software that generates language by predicting patterns across enormous volumes of text, misreads laboratory results often enough that no one should hand it an open-ended clinical question. We build a lab interpretation report that runs on this same class of software. So we will show you the studies first, and then we will show you exactly what we did about them.

Where these systems fail

In 2023, a team of clinical biochemists at Gloucestershire Hospitals in England gave fifteen fictional thyroid cases to two publicly available chatbots and graded the answers against a consensus opinion from practicing clinicians and clinical scientists. ChatGPT interpreted 33.3 percent of the cases correctly. Google Bard managed 20.0 percent. Both programs recognized straightforward overactive and underactive thyroid states. Both programs failed on subclinical presentations, on illness elsewhere in the body that shifts thyroid numbers while the gland itself works fine, and on secondary hypothyroidism, where the failure sits in the pituitary gland above the thyroid. The authors concluded that these tools cannot replace human clinical knowledge in this setting.

A second study appeared that same year in Annals of Clinical and Laboratory Science. Researchers at Baylor College of Medicine put 35 real clinical chemistry questions to two versions of the same technology and to residents, fellows, and faculty. Senior chemistry faculty answered 100 percent correctly. The older chatbot reached 60 percent. The newer one reached 71.4 percent, which matched the score of first-year residents.

A third study, in Radiology in 2024, gave the vision-capable version of the same model 72 case questions from the Radiological Society of North America annual meeting. It scored 43 percent. When the researchers removed the written case history and left only the images, the score fell to 38 percent. The model had been working from the text the whole time.

Three studies, one pattern. Open-ended clinical reasoning breaks these systems.

Where these systems succeed

Now the other half of the evidence, which almost nobody quotes.

In 2025, anesthesiologists in Istanbul ran 400 arterial blood gas samples from intensive care patients through the newer model and had two trained physicians grade every output. The model reached 100 percent accuracy on acidity, oxygenation, sodium, and chloride. It reached 92.5 percent on hemoglobin. It fell to 72.5 percent on bilirubin, the pigment the liver clears from blood, and it sometimes recommended bicarbonate treatment for patients who had no need of it.

An emergency physician in Istanbul repeated the exercise later that year across 25 blood gas scenarios. Agreement with the specialist ran above 90 percent for chronic obstructive pulmonary disease, asthma, and fluid in the lungs. Agreement fell below 70 percent for poisonings and for mixed acid-base states carrying two competing problems at once.

Set those numbers beside the thyroid study. The same technology reaches near-perfect accuracy on one kind of task and roughly one answer in three on another.

The line between the two

The dividing line holds across every study above, and it has nothing to do with how difficult the medicine is.

These systems perform when three conditions hold together. The input arrives as structured numbers. The reference values sit fixed and defined in advance. The question asks for a comparison.

These systems fail when the question has many possible answers. Which of forty diagnoses fits this patient. What does this shadow on the film represent. Which of three competing metabolic problems came first. Every failure listed above sits in that second category.

That distinction decided how we built our report.

How our report is built

Our report lives entirely in the first category, by design.

You upload a standard blood panel. The system reads each marker as a number. It compares that number against a reference range a physician defined in advance, drawn from the published functional and optimal range literature and corrected for the population the marker belongs to. It returns which markers sit inside that range, which sit outside, and by how much.

That output drives a written plan covering the foods to prioritize, the supplement priorities your specific markers call for, when to eat, and how to sleep and move. Every recommendation in that plan traces back to a named marker and a defined range. The system proposes no diagnosis. The system chooses among no competing conditions. The system reads no images.

A physician designed the ranges. A physician built and reviews the report logic. The scope stays fixed.

What this report will never do

It will never diagnose you. It will never prescribe for you. It will never stand in for the clinician who examines you, knows your history, and carries the license to treat you.

Everything it returns is educational. Bring it to your physician.

Sources

  1. Stevenson E, Walsh C, Hibberd L. Can artificial intelligence replace biochemists? A study comparing interpretation of thyroid function test results by ChatGPT and Google Bard to practising biochemists. Annals of Clinical Biochemistry. 2023;61(2):143-149. DOI
  2. Ibrahim RB, Chokkalla AK, Levett K, Gustafson D, Olayinka L, Kumar S, Devaraj S. ChatGPT: Exploring Its Role in Clinical Chemistry. Annals of Clinical and Laboratory Science. 2023;53(6):835-839. PubMed 38182139
  3. Mukherjee P, Hou B, Suri A, et al. Evaluation of GPT Large Language Model Performance on RSNA 2023 Case of the Day Questions. Radiology. 2024;313(1):e240609. DOI
  4. Turan Eİ, Baydemir AE, Balıtatlı AB, Şahin AS. Assessing the accuracy of ChatGPT in interpreting blood gas analysis results. Journal of Clinical Anesthesia. 2025;102:111787. DOI
  5. Gün M. AI-Assisted Blood Gas Interpretation: A Comparative Study With an Emergency Physician. American Journal of Emergency Medicine. 2025;94:1-2. DOI

What the report includes

Important Disclaimers

This content is for informational purposes only and is not intended as medical advice, diagnosis, or treatment. Always consult a qualified healthcare professional before making changes to your diet or health routine.

These statements have not been evaluated by the Food and Drug Administration. No product or service offered here is intended to diagnose, treat, cure, or prevent any disease. Nothing on this platform is a clinical genetic test and no laboratory testing is performed by Geneflammation.

Decoding Inflammation

The GENEFLAMMATION platform — physician-designed precision wellness education by Alberto Rivera, MD, FAAPMR.

HIPAA-Conscious. Your health data is always private and protected.

Contact

Dr. Rivera practices medicine in the state of Florida only. This platform does not establish a physician-patient relationship.

This content is for informational purposes only and is not intended as medical advice, diagnosis, or treatment. Always consult a qualified healthcare professional before making changes to your diet or health routine.

© 2025 Decoding Inflammation / GENEFLAMMATION. All rights reserved.