LIVE
1.29°S / 36.82°E  ·  Nairobi

Estimating the fatality rate is difficult but doable with better data

Editorial note, September 3, 2026. This article preserves an argument published on August 3, 2020, during the acute phase of the COVID-19 pandemic. Its terminology has been clarified rather than silently modernized: the proportion of deaths among identified cases is the case fatality ratio (CFR), while the proportion among all infected people is the infection fatality ratio (IFR). The original article and paper used CFR for the latter quantity. On May 5, 2023, the World Health Organization determined that COVID-19 no longer constituted a public health emergency of international concern, while describing it as an established and ongoing health issue.

Fatality ratios help describe disease severity and how the risk of death varies across populations. In 2020, however, estimates based on reported COVID-19 cases and deaths were vulnerable to under-ascertainment, outcome delays, changing case definitions, incomplete reporting, and non-random testing. Our central claim was simple: more observations cannot repair systematic selection bias. Better sampling can.

What are we trying to estimate?

A CFR estimates the proportion of identified cases that end in death. An IFR instead uses all infections, including infections never diagnosed. Both quantities depend on the population, period, outcome definition, and data-generating process. Relative ratios between subpopulations can help describe differences in risk, but they do not by themselves determine how resources should be allocated.

Suppose we want the IFR but calculate reported deaths divided by reported cases. We call this the naive estimator, E_naive. Its numerator and denominator are not merely uncertain: they can be selected in systematically different ways. Mild or asymptomatic infections may be missed, while severe infections are more likely to reach clinical attention. During a growing outbreak, recently diagnosed cases may not yet have resolved, so deaths observed today belong partly to infections diagnosed earlier.

Why raw fatality ratios can be biased

In the United States during early 2020, jurisdictions supplied aggregate counts and voluntarily submitted individual case reports to the Centers for Disease Control and Prevention. The CDC cautioned that this surveillance captured only a subset of infections, contained substantial missing data, and represented asymptomatic infections poorly. The WHO likewise noted that limited testing could concentrate detection among severe cases and priority groups, while unresolved cases created downward pressure on crude CFR calculations during a growing epidemic.

A related illustration came from Xiao-Li Meng’s data-defect-correlation framework. Under an assumed correlation of 0.005 between being tested and testing positive, selectively testing 10,000 people for SARS-CoV-2 prevalence was calculated to carry roughly the information of a random sample of 20 people—a 99.8% reduction in effective sample size. That calculation concerned prevalence, not fatality directly, but it illustrates how a small selection correlation can overwhelm a large nominal sample.

“Compensating for data quality with quantity is a doomed game.” — Xiao-Li Meng

The original paper grouped the principal problems into five broad categories: under-ascertainment of mild infections, time lags, interventions, population characteristics, and imperfect reporting or attribution. Its graphical model also made an identification point: relationships outside the surveillance sampling frame cannot generally be recovered from aggregate observations inside that frame without additional assumptions or data.

Biases can also cancel. Suppose under-ascertainment makes the naive estimator too high by a factor b, while unresolved outcomes make it too low by the same factor. The two errors could produce the correct number for the wrong reasons. Correcting only the outcome lag would then move the estimate farther from the target. This hypothetical example does not argue against correction; it argues for stating the complete model and testing its assumptions.

What the mathematical model says

In the simplified model developed in our Harvard Data Science Review paper, let p be the infection fatality probability, q the probability that an infection is reported, r the covariance between eventual death and reporting, and N the number of infected individuals considered. Under the paper’s conventions, the expectation of the naive estimator is:

μ = (r/q + p)(1 − (1 − q)^N).

As N grows, the final factor approaches one, but the selection term r/q remains. If fatal infections are more likely to be reported, r is positive and the limiting estimate exceeds p. More observations drawn through the same mechanism reduce variance but do not remove that term. Under this model, asymptotic unbiasedness requires r = 0; exact finite-sample unbiasedness, under the paper’s treatment of samples with no reported infections, requires perfect reporting.

This is the durable point: selection bias does not disappear merely because the dataset becomes large.

A prospective contact-based strategy

We proposed collecting a smaller but more deliberate dataset through contact investigation. The reconstructed protocol is:

1. Identify an infected index case through ordinary clinical or surveillance procedures.

2. Recruit exposed contacts before their eventual symptom severity and outcome are known, asking them to agree to testing and follow-up.

3. Test at a prespecified time appropriate to the exposure and assay, regardless of symptoms.

4. Retain granular individual-level records subject to consent, privacy, ethics, and law.

5. Follow participants long enough to ascertain resolved outcomes rather than comparing current diagnoses with current deaths.

6. Record nonresponse, missed tests, loss to follow-up, and reported symptoms. A symptom report alone does not establish whether infection occurred.

This design can reduce the association between severity and diagnosis because recruitment occurs before severity is known. It does not guarantee independence. Participation and follow-up can still differ by health, access, behavior, or eventual severity; traced contacts form a restricted sample rather than an automatic random sample of the population; and imperfect tests remain a source of error. The primary paper itself therefore described contact tracing as a way to limit diagnosis–death covariance, not eliminate every bias.

What the sample-size calculation does—and does not—show

When r = 0, the paper derived a bound for making the estimator’s expected value fall within δ of p:

N ≥ log(δ/p) / log(1 − q).

The paper’s table reported N = 66 when the reporting probability was q = 0.1 and the permitted relative bias was δ/p = 0.001. The number 66 therefore bounds bias in the estimator’s expectation under an idealized model. It does not imply that 66 observations can estimate a fatality probability of 0.001 precisely. When death is rare, observing enough deaths to control variance requires a much larger sample—on the order of 1/p merely to expect one death.

Even a study with no observed deaths contains information. If 1,000 independently sampled infected people were followed completely and none died, a true fatality probability of 1% would yield that result with probability about 0.000043. Under the same assumptions, the one-sided exact 95% upper confidence bound would be about 0.3%, not 1%. Such statements cease to hold if infections are misclassified, outcomes are missing, observations are dependent, or the sample is not representative of the target population.

From a traced sample to a target population

An estimate valid for traced contacts is not automatically valid for a city, country, age group, or other target population. Post-stratification can adjust for measured differences when the population composition of relevant strata is known. Multilevel regression and post-stratification can stabilize estimates across many cells, but it cannot correct unmeasured selection or guarantee that sampled and unsampled people are comparable within each cell.

Our academic paper also adapted a likelihood model developed by Reich and colleagues to address outcome delays and relative reporting rates. Such models can improve estimates when their assumptions are approximately correct. They cannot identify information absent from the data-generating process.

CFR and IFR estimation was—and remains—a problem of defining the estimand, understanding how observations entered the dataset, following outcomes to resolution, and reporting uncertainty. Sophisticated estimation is useful, but it comes after the more basic work of constructing data that bear a defensible relationship to the population and question of interest.

Source. This restoration is based on the Berkeley Artificial Intelligence Research Blog post by Anastasios Angelopoulos, Reese Pathak, and Michael I. Jordan, published August 3, 2020, and on On Identifying and Mitigating Bias in the Estimation of the COVID-19 Case Fatality Rate by Anastasios Nikolas Angelopoulos, Reese Pathak, Rohit Varma, and Michael I. Jordan, published in 2020.

Responses