Back to courses

LESSON 90 · Environment and public health

How epidemiology studies disease in populations

Case counts become informative through denominators, time, and credible comparison. Hypothetical data illustrate prevalence, incidence risk, and association measures, while study design, bias, confounding, and causal reasoning explain what conclusions are justified.

What you will be able to do

  • Distinguish prevalence, incidence risk, person-time rates, and association measures.
  • Identify major study designs and errors through sampling and comparison.
  • Explain differences among statistical uncertainty, population causal evidence, and individual attribution.
In this lessonDefine the population, outcome, and denominatorPrevalence, risk, and incidence rate answer different questionsComparison groups shape the question a study can answerExpress relative and absolute differences togetherSelection, measurement, and confounding can distort associationSeparate statistical uncertainty from causal evidenceUse evidence for action while continuing evaluationBilingual termsSources

Define the population, outcome, and denominator

Epidemiology studies the distribution of health events, factors shaping that distribution, and the application of evidence to prevention and control. It includes injuries, chronic illness, function, and service outcomes as well as infections. Clinical work often asks why one person feels unwell; epidemiology also asks which people experience an outcome under which conditions. These questions can inform each other, but individual diagnosis and population comparison require different information and levels of inference. CDC textbook: Definition of epidemiology

Before comparison, specify who can enter the study, what counts as a case, and the observation period. A case definition may combine symptoms, testing, place, and time so investigators count consistently; it need not contain every requirement of a clinical diagnosis. More cases in one area may reflect a larger population or more testing. Counts need corresponding denominators and examination by time, place, and population characteristics before an apparent difference can be interpreted. CDC textbook: Descriptive epidemiology CDC textbook: Steps of an outbreak investigation

Prevalence, risk, and incidence rate answer different questions

Prevalence describes the proportion living with a condition at a point or during a period, including earlier cases still present. Incidence risk concerns new cases among people initially free of the target outcome and capable of developing it during a specified period. In a hypothetical population of 1,000 with 100 existing cases, point prevalence is 10%. If the other 900 complete one year of follow-up and 45 develop disease, their one-year risk is 5%. The numerators, denominators, and time meanings differ. CDC textbook: Measures of morbidity

With unequal follow-up, investigators can sum observed time during which each person remains at risk. New cases divided by this person-time form an incidence rate, expressed, for example, as cases per 1,000 person-years. This describes occurrence speed and is not directly each person’s future annual probability. Prevalence also depends on duration and survival: longer survival can increase existing cases without more new disease. Choosing a measure should follow the prevention or service question being asked. CDC textbook: Measures of morbidity

Hypothetical complete two-year cohort: measures are not interchangeable

MeasureCalculation and unitInterpretation
Exposed risk20/200 = 10%, over two yearsProportion developing disease in the exposed group
Unexposed risk10/200 = 5%, over two yearsComparator proportion over the same period
Risk ratio10%/5% = 2, unitlessTwice the comparator risk
Risk difference10% − 5% = 5 percentage pointsFive additional cases per 100 over two years
Incidence rateRequires actual person-time dataCannot be calculated exactly from these counts alone

Comparison groups shape the question a study can answer

Cross-sectional studies measure exposure and health status during a period and describe existing burden, often without establishing order. Cohorts organize people by exposure or another starting point and compare later outcomes, prospectively or by reconstructing existing records. Case-control studies start with cases and appropriate controls and compare earlier exposures, often efficiently studying rare outcomes. The key distinction is how people were selected and whether controls represent the source population that produced the cases, not simply whether questions concern the past. CDC textbook: Analytic epidemiology

Randomized trials allocate interventions and reduce systematic effects of pre-existing group differences, but missing follow-up, implementation, and outcome measurement still matter. Assigning people to smoke or encounter known hazards would be unethical, so many public health questions require observational evidence, comparison across designs, and mechanisms. No design is best for every question. Its value depends on credible comparison, ethical conduct, and relevance to the target population and real service conditions. CDC textbook: Analytic epidemiology

Express relative and absolute differences together

Suppose a hypothetical two-year cohort includes 200 exposed and 200 unexposed participants, all initially disease-free with complete follow-up. Twenty exposed and ten unexposed people develop disease. Risks are 10% and 5%, the risk ratio is 2, and the risk difference is five percentage points. The ratio describes change relative to the comparator; the difference describes additional cases per 100 over the period. Saying only that risk doubled does not establish whether the outcome is common or rare. CDC textbook: Measures of association

An odds ratio compares odds rather than probabilities. A probability of one in ten corresponds to one event for nine nonevents. Case-control studies commonly estimate odds ratios, but the proportion of cases selected by investigators is not population risk. Under appropriate sampling conditions, an odds ratio for a rare outcome may approximate a risk ratio; for common outcomes they can differ substantially. Preserve the measure’s name, comparator, and period rather than translating every multiplicative estimate into the same personal risk. CDC textbook: Measures of association

Selection, measurement, and confounding can distort association

Selection bias occurs when entry or retention distorts the target comparison, for example if ill participants with higher exposure are more likely to disappear from follow-up. Information bias involves systematic measurement errors, such as cases recalling past diet more carefully than controls. A larger sample generally reduces random variation without automatically repairing these problems. Collecting more of the same flawed measurements may make a biased result appear precise. Recruitment, missingness, measurement procedures, and plausible directions of error need examination. CDC: Analyzing and interpreting data

Confounding mixes the influence of another factor with the exposure under study. If workers in one occupation are older, and age affects disease risk, part of the crude occupational association may reflect age composition. Comparison within age strata or statistical adjustment under stated assumptions can help. Extensive adjustment does not guarantee causal truth: omitted or mismeasured factors and inappropriate adjustment for processes caused by the exposure can change the interpretation. The question and causal structure should guide the choice of variables. CDC: Analyzing and interpreting data

Separate statistical uncertainty from causal evidence

A confidence interval expresses estimation uncertainty under a statistical model and assumptions; a wider interval usually indicates lower precision. A 95% confidence level describes long-run coverage when the method is repeatedly used. It does not mean a result has a 95% probability of being unbiased or that an individual has a 95% probability of falling inside it. Significance does not establish effect size or public health importance: large studies may detect tiny differences, while small studies may miss important ones. CDC: Confidence intervals

Causal assessment also examines temporal order, alternative explanations, consistency, mechanisms, and intervention evidence. Many diseases arise through combinations of conditions. A factor can contribute without guaranteeing illness in every exposed person or being necessary for every case. Even when population evidence supports preventing disease by reducing an exposure, one person’s illness does not establish that the exposure caused that particular case. The causal question about population prevention differs from attribution of an individual illness. CDC textbook: Concepts of disease occurrence

Use evidence for action while continuing evaluation

Outbreak investigation combines case definitions, active case finding, patterns by time and place, comparison studies, laboratory findings, and environmental evidence. When sufficient information indicates continuing harm from a source, suitable control can proceed alongside investigation without waiting for every mechanistic detail. New cases and intervention burdens still need monitoring. A decline must be interpreted alongside natural epidemic dynamics, testing changes, and other measures rather than attributing the entire change to whichever action occurred first. CDC textbook: Steps of an outbreak investigation

For any health statistic, ask who was studied, what was measured, what was compared, where error might arise, and where conclusions apply. Do not automatically convert an area-level association into an individual relationship: high average exposure and more disease in an area do not establish the exposure experienced by each ill resident. Maps and grouped data generate useful questions, while individual judgments require relevant individual information. Sound epidemiology supports action while keeping unresolved uncertainty visible. CDC textbook: Definition of epidemiology CDC textbook: Descriptive epidemiology

Apply what you have learned

In a hypothetical complete two-year cohort, 20 of 200 exposed and 10 of 200 unexposed people develop disease. Calculate risks, risk ratio, and risk difference. What is needed if exposed participants are substantially older, and why can an exposed individual case not automatically be attributed to that exposure?

Read the explanation

Two-year risks are 10% and 5%, the risk ratio is 2, and the risk difference is five percentage points, or five additional cases per 100 over two years. Examine age strata and other plausible confounders, state adjustment assumptions, and assess selection and measurement errors. Association is not yet causation; even stronger population causal evidence does not reveal this person’s outcome without the exposure and cannot establish certain individual attribution.

Bilingual terms

患病率 · prevalence
Proportion with a target condition at a specified point or during a period.
发病风险 · incidence risk
Proportion of an initially at-risk population developing a new outcome over a specified period.
人时 · person-time
Sum of time people are observed and at risk of the target outcome.
风险比 · risk ratio
Risk in the exposed group divided by risk in the comparator.
混杂 · confounding
Distortion of causal interpretation of an exposure–outcome association by other factors.
选择偏倚 · selection bias
Systematic distortion of the target comparison through selection or retention.
置信区间 · confidence interval
An interval expressing estimation uncertainty under specified statistical assumptions.

Sources and further reading

Original course source-check record: 9 September 2026. Full Chinese and English sentence-by-sentence language review: 14 September 2026. AI editing and language review are not human clinical review. Linked institutions have not participated in or endorsed this course.

A moment in natureLayers of orange and purple dunes in the Namib Desert.

Namib-Naukluft Sand Dunes (2011).jpg · Yathin S Krishnappa · CC BY-SA 3.0
Converted to WebP; thumbnails may be cropped.