LESSON 78 · Disease, medicine and care
Interpreting tests: reference ranges, error and false positives
Test reports compress complex physiology into numbers or categories. Interpretation requires measurement context, reference ranges, disease probability, and an understanding of possible errors.
What you will be able to do
- Distinguish reference ranges, diagnostic thresholds, and treatment goals.
- Explain sensitivity, specificity, and positive predictive value using counts.
- Explain how repeat testing and context can change interpretation.
In this lesson
From sample to reported numberA reference interval does not separate health from diseaseWhat sensitivity and specificity measureUse ten thousand people to interpret a positiveWhy repeat testing may answer a new questionReasoning: what should be checked first?Bilingual termsSourcesFrom sample to reported number
A result depends on collection, transport, analysis, and reporting. Food, activity, medicines, and sampling time can influence some measurements, and sample handling may introduce bias. Even quality-controlled instruments have measurement uncertainty, while the body varies within and between days. A difference between yesterday and today therefore does not automatically prove deterioration or laboratory failure. First ask whether the conditions are comparable.
Preparation should follow the instructions for the particular test. Not every blood test requires fasting, and medicines should not be stopped independently to obtain a more attractive result. Tell staff if preparation instructions were not followed so they can decide whether to proceed and how to interpret the sample. Record the test, units, reference interval, and date. Laboratories may use different methods or units; ignoring these details can turn a reporting difference into an apparent dramatic biological change.
A reference interval does not separate health from disease
Many reference intervals describe a distribution in a selected reference population, often its central approximately 95%, although not every test uses that method. Healthy people can fall outside the interval, and people with disease can fall within it. Age, pregnancy, and other factors may change the applicable interval. An arrow on a report indicates a difference from the laboratory interval; it does not independently identify a disease or determine severity.
Diagnostic thresholds reflect disease definitions and evidence. Treatment targets incorporate expected benefits, risks, and personal context. These three kinds of numbers may appear similar but have different purposes. A result just above an upper limit in someone without symptoms calls for attention to magnitude, history, and risk. Severe symptoms in someone with an in-range result still warrant assessment. Reference intervals organize measurement; they cannot replace judgment about the person.
Key distinctions and reasoning cues
| Concept or situation | Meaning | Reasoning focus |
|---|---|---|
| True positive | Disease present, test positive | 90 in the example |
| False negative | Disease present, test negative | 10 in the example |
| False positive | Disease absent, test positive | 495 in the example |
| True negative | Disease absent, test negative | 9,405 in the example |
What does a positive result imply?
Assume 10,000 people are tested. All inputs are invented. Outputs are expected counts and may include fractions. These figures do not represent the performance of an actual test. Hold sensitivity and specificity constant and lower prevalence to examine positive predictive value. In practice, performance can also vary with population and conditions.
What sensitivity and specificity measure
Test performance can be assessed against an appropriate reference standard. Sensitivity asks what proportion of people who truly have the target condition test positive. Specificity asks what proportion without it test negative. A false negative misses the target condition; a false positive reports it when it is absent. Performance may vary with disease stage, threshold, sample quality, and study population, rather than being an eternally fixed property of a machine.
Making the positivity threshold easier to cross often reduces missed cases while increasing false alarms, with the trade-off determined by the data. Evaluation should also ask whether the reference standard is reliable, participants receive comparable verification, and the study resembles intended use. High sensitivity does not mean every positive is true. High specificity does not mean a negative result explains every symptom. Questions about an individual result also require the probability before testing.
Clinical Methods: Sensitivity, Specificity, and Predictive Value
Use ten thousand people to interpret a positive
This is an invented teaching example, not the performance of a real test. Among 10,000 people, suppose 100 have a condition. Sensitivity is 90% and specificity is 95%. There are 90 true positives and 10 false negatives. Among the 9,900 people without the condition, there are 9,405 true negatives and 495 false positives. Of 585 positive results, only 90 are true positives: positive predictive value is about 15.4%.
There is no contradiction. A small false-positive fraction applied to a much larger disease-free group can outnumber detected cases. With a higher prevalence and unchanged performance assumptions, the true-positive share of positive results would rise. This conditional-probability calculation does not make all low-prevalence testing useless; consequences of missed disease, confirmation, and treatment benefit still matter. Translating 90% sensitivity into a 90% personal chance of disease after a positive result changes the denominator incorrectly.
Clinical Methods: Sensitivity, Specificity, and Predictive Value
Why repeat testing may answer a new question
Repeat testing may assess temporary variation, improve collection conditions, or observe disease at a more informative time. A confirmatory test sometimes uses a different principle to help distinguish signal from a false alarm. Repeating an identical measurement does not automatically remove systematic bias: persistent interference can produce consistently misleading results. The purpose of repetition should specify which uncertainty it is intended to resolve.
Ordering many weakly relevant tests also increases the chance of at least one out-of-range result. If 20 independent tests each have a 5% chance of such a finding, the probability of at least one is 1 minus 0.95 to the twentieth power, approximately 64%. Real tests are often correlated, so this number cannot be applied mechanically. The example explains why more testing does not automatically mean more reliability. Selection should answer a defined question, and follow-up should depend on clinical meaning rather than arrows alone.
Reasoning: what should be checked first?
Wu sees three mildly abnormal results and concludes from internet searches that three serious diseases are present. First check the tests, units, applicable intervals, preparation, and previous results, then discuss symptoms and risk. The number of flags cannot measure danger, and a small deviation does not guarantee that no action is needed. Some findings merit repetition, others further evaluation, with the plan determined by context.
If severe new chest pain, altered consciousness, or another dangerous symptom is present, obtaining help takes priority over interpreting the report. Without immediate danger, useful visit questions are why the test was done, what probabilities it changed, and whether the next step would change treatment. The goal is not memorizing every number. It is understanding what results support, what they cannot establish, and who will complete the next action and when. That converts measurement into a potentially useful decision.
Apply what you have learned
In the example, 585 test positive and 90 truly have disease. What is positive predictive value, and why does it differ from sensitivity?
Read the explanation
90/585 is about 15.4%. Positive predictive value uses all positive results as the denominator; sensitivity uses all truly affected people, here 90/100 = 90%.
Bilingual terms
- 参考范围 · Reference interval
- An interval based on measurements in a specified reference population.
- 敏感度 · Sensitivity
- Proportion of truly affected people who test positive.
- 特异度 · Specificity
- Proportion of unaffected people who test negative.
- 假阳性 · False positive
- A positive result when the target condition is absent.
- 阳性预测值 · Positive predictive value
- Proportion of positive results belonging to truly affected people.
Sources and further reading
- MedlinePlus: How to Understand Your Lab Results
- NCI: Cancer Screening Overview
- MedlinePlus: How to Prepare for a Lab Test
- Clinical Methods: Sensitivity, Specificity, and Predictive Value
- Clinical Methods: Use of the Laboratory
- National Academies: The Diagnostic Process
Original course source-check record: 9 September 2026. Full Chinese and English sentence-by-sentence language review: 14 September 2026. AI editing and language review are not human clinical review. Linked institutions have not participated in or endorsed this course.
A moment in nature

Sunlight in the forest.jpg · Antoloji · CC BY-SA 4.0
Converted to WebP; thumbnails may be cropped.
