LESSON 07 · Body structure and function
How Health Research Works: Experiments, Observation, and Causation
A study finds that regular walkers are healthier. Does walking make the difference, or are healthier people more able to walk regularly? Good research uses careful comparisons to distinguish these explanations.
What you will be able to do
- Turn a broad health claim into a clear research question.
- Distinguish observational studies from randomized trials and explain confounding, bias, and masking.
- Assess a conclusion using its population, outcomes, conduct, and ethical limits.
In this lesson
Start with a question that can be answeredWhat observational studies can revealCheck whether the groups are comparableRandom allocation and masking address different problemsMatch the evidence to the questionRead a finding in the context of other evidenceBilingual termsSourcesStart with a question that can be answered
“Walking is good for health” can lead to many different studies. One might ask whether starting a walking program changes blood pressure in sedentary adults. Another might examine falls among older people. The populations, comparisons, time frames, and outcomes differ, so the findings are not interchangeable. A useful question identifies who is being studied, what exposure or intervention matters, what it will be compared with, and how change will be measured.
Consider an invented example: researchers offer sedentary adults twelve weeks of walking support and compare their walking endurance with that of adults given written health information. The program may include reminders, group activities, and guidance. A difference would concern this package; it would not establish which component produced it. Researchers also need to specify how endurance will be measured and what improvement would matter in everyday life.
Ask what result would change your mind. If both groups improved equally, would you accept that the support program had shown no additional advantage? Measuring change before and after participation alone cannot separate program effects from improvement through familiarity with the test. Thinking through alternative results helps prevent a search for evidence that merely confirms a preferred answer.
What observational studies can reveal
Observational studies record differences that occur in life without researchers assigning the exposure. A cohort study can record activity habits and follow subsequent health. A case-control study starts with people who do or do not have a particular condition and compares earlier exposures. A cross-sectional survey measures exposure and health at roughly the same time; it can describe a population but often cannot establish which came first.
If regular walkers experience less illness, walking is one possible explanation. Age, existing disease, smoking, income, and neighborhood conditions may also matter. Some people may reduce activity because illness has already begun: this is reverse causation. Recording exposure before the outcome helps establish sequence, but does not remove every alternative explanation. Observational research remains essential for long-term exposures, uncommon outcomes, and harmful conditions that could not ethically be assigned.
Explore the concept
Hypothetical study: people with more time to exercise sleep longer. The diagram proposes one possible confounder: work schedules affect both. These arrows are a hypothesis, not established causation; the study needs relevant measurements and alternative explanations.
Check whether the groups are comparable
Confounding is a major obstacle to causal interpretation. Suppose the walkers are younger, and age also affects the outcome. The difference between groups then includes an age-related difference. Researchers may restrict the age range, compare people within age groups, or adjust measured characteristics in a statistical model. Unmeasured factors and imperfect measurements can still leave residual confounding.
Bias can also enter through recruitment, measurement, or follow-up. Recruiting only enthusiastic volunteers may leave out people who find activity hardest. Measuring walking with a device in one group and memory in another creates an uneven comparison. Missing follow-up matters too: if people leave because the program causes difficulties, ignoring them may make the result look better. A larger sample can reduce random fluctuation, but it cannot automatically repair a systematically distorted comparison. Participant numbers should therefore prompt questions about methods, not end them.
Random allocation and masking address different problems
In a randomized trial, chance determines which program a participant is offered. The aim is to balance starting characteristics, including unknown ones, rather than guarantee identical groups in every trial. The next allocation should also be concealed from recruiters so that knowledge of it cannot influence enrollment.
Masking addresses a different problem: knowing the treatment can change expectations, behavior, and assessment. Participants in a walking trial cannot realistically be unaware that they are walking. However, the assessor who measures endurance may remain unaware of group assignment, and both groups can undergo the same testing procedure. Lack of double masking does not make a study worthless; it identifies a limitation to examine. Researchers must also explain withdrawals, treatment switching, and missing data. Selecting only the most successful participants after randomization can undermine the comparability that random allocation was intended to create.
Match the evidence to the question
Laboratory work can reveal cellular mechanisms, and animal experiments can explore responses in an organism. Human doses, metabolism, behavior, and disease contexts may differ. Changing cells in a dish does not establish that a product improves human health. Human studies also have boundaries: improved endurance after twelve weeks does not demonstrate longer life, and a blood marker is not automatically equivalent to relief of symptoms or fewer hospital admissions.
Ethics also determines which questions can be tested experimentally. Researchers cannot deliberately expose people to harmful air for years simply to isolate its effects. Ethical research requires independent scrutiny, meaningful consent, proportionate risks, and respect for withdrawal. Randomized trials are valuable for many intervention questions, but they need support from other methods when studying rare harms, long-term safety, implementation in ordinary settings, and the experiences of people receiving care.
Read a finding in the context of other evidence
When a headline announces a new discovery, find the underlying report and identify the actual question. How do participants differ from you? Was the highlighted result a prespecified main outcome, or one selected from many analyses? Was the study registered, were outcomes reported fully, and were funding and competing interests disclosed? These questions identify potential exaggeration more effectively than relying on an author's reputation.
Look across independent studies as well. A systematic review uses explicit methods to find and assess relevant research; a meta-analysis combines numerical results when appropriate. Pooling fundamentally different participants, interventions, or outcomes does not automatically produce a trustworthy answer. Try summarizing the evidence with its limits intact: this study supports a particular possible benefit, in a defined population, over a stated period, with specified uncertainties. Understanding what a result cannot establish is part of understanding what it can.
Apply what you have learned
A walking class advertises that the 40 people who completed twelve weeks improved their average endurance, proving that the course suits everyone and extends life. Identify at least four missing details, then rewrite the claim to match the information available.
Read the explanation
We need enrollment and withdrawal numbers, a suitable comparison group, baseline health, measurement methods, effect size and uncertainty, and information about adverse experiences. The report describes change among completers; it does not isolate the cause or measure lifespan. A proportionate claim is: these 40 completers improved on average on an endurance measure after twelve weeks; whether the program caused the change, who else might benefit, and its long-term effects require further study.
Bilingual terms
- 混杂 · Confounding
- The influence of other factors becomes mixed with the exposure–outcome relationship, complicating causal interpretation.
- 反向因果 · Reverse causation
- The presumed outcome influences the behavior or exposure being studied.
- 盲法 · Masking
- Keeping specified participants or researchers unaware of allocation to reduce bias from expectations, behavior, or assessment.
Sources and further reading
- NIH:临床试验基础
- CDC:现场分析研究的设计与实施
- CDC:资料分析与解释
- CDC:资料收集与误差
- NCCIH:阅读健康新闻的检查问题
- Cochrane:解释研究结果与形成结论
- Cochrane:合并分析及其局限
Original course source-check record: 9 September 2026. Full Chinese and English sentence-by-sentence language review: 14 September 2026. AI editing and language review are not human clinical review. Linked institutions have not participated in or endorsed this course.
A moment in nature

Khrangsuri waterfall, Meghalaya 01 (edit).jpg · Original: Chirnzb Derivative work: UnpetitproleX · CC BY-SA 4.0
Converted to WebP; thumbnails may be cropped.
