Research with EHR data - Pitfalls

Read a paper and hunt for these pitfalls. Mark a square only when you can point to the sentence that shows it.

Glossary

Coding & vocabulary
EXPOSURE DEFINITION
The rule used to decide who counts as exposed. Too broad or too narrow, and it misclassifies who actually got the exposure.
CONCEPT SET
The list of codes standing in for one clinical idea, often built from a parent code plus its descendants. Pulling descendants in bulk without reviewing them sweeps in codes that do not belong.
VOCABULARY MAPPING
Translating codes from one coding system into a different standard vocabulary; distinct clinical meanings can collapse into one broader code.
CODE AVAILABILITY
Whether a code existed and was in use for the whole study period. A code introduced partway through can't capture anything before it existed.
SITE HETEROGENEITY
Different sites coding the same clinical event differently. Pooling sites without accounting for this can mistake coding differences for differences in disease.
Who's in the data
SELECTION BIAS
Bias from how the cohort was assembled: among the people selected, the exposure-outcome relationship differs from the one in the population they stand in for. It arises when entry depends on something tied to both the exposure and the outcome, such as reaching a particular care setting.
COMPARATOR CHOICE
Who or what the exposed group is measured against. An unrelated or non-comparable comparator can't isolate the exposure's effect.
LEFT TRUNCATION
Observation begins only when a person or site enters the data source, not when the clinical story started. Anyone whose event happened before that date never appears at all, and for those who do appear, earlier history is missing.
CARE FRAGMENTATION
Care received outside the data source, such as a different health system or unlinked records, that the database simply can't see.
SURVEILLANCE BIAS
Outcomes are detected and coded more often in people with more healthcare contact. When one group is watched more closely than the other, it looks like more disease when it may be only more looking.
Rigor & reporting
PHENOTYPE VALIDATION
Formally checking how well a code-based definition identifies true cases (sensitivity, specificity, PPV) against a reference standard.
UNAPPLIED STANDARD
A methodological threshold or rule the study states it will use but never actually applies to its own results.
CITATION ACCURACY
Whether a cited reference actually supports the specific claim it's attached to.
Time & follow-up
TIME ZERO
The point where follow-up starts, which should also be where eligibility is met and treatment is assigned. Anchoring the two arms to different kinds of events makes them non-comparable from the first day.
IMMORTAL TIME BIAS
A stretch of follow-up counted as exposed even though anyone who had the outcome during it could never have been classified as exposed. Because events in that window cannot land in the exposed group, the treatment looks more protective than it is.
PERSON-TIME
The time each person is actually under observation and at risk, which is the denominator of an event rate. Reducing years of follow-up to a yes/no lets someone followed six months count the same as someone followed five years.
CENSORING
Follow-up ending before the outcome is seen, whether the person leaves the system, switches care, or the study period closes. Ignoring it treats "unobserved" as "unaffected."
REVERSE CAUSATION
The outcome, or its early signs, actually preceded and possibly caused the exposure, rather than the other way around.
Confounding & causation
UNMEASURED CONFOUNDING
A variable that affects both the exposure and the outcome but is not captured in the data, so no amount of adjustment can remove its effect on the estimate.
CONFOUNDING BY INDICATION
Treatment is chosen because of the patient's underlying condition or prognosis, so the treated differ from the untreated before anyone is treated. Matching on demographics does not touch it, because the reason for treatment is itself a risk factor for the outcome.
CRUDE ESTIMATE
An association reported without adjusting for confounders: the raw, unadjusted comparison.
CAUSAL LANGUAGE
Wording that asserts X causes Y when the design and its stated assumptions cannot support it. Observational data can support causal claims, but only when the design is built for that and the assumptions are named.
ADHERENCE
Whether a patient actually took or used a treatment as prescribed. Assuming a prescription equals full adherence overstates true exposure.