Genetic Prediction and Adverse Selection
with Richard Karlsson Linnér and Jonathan Beauchamp
Introduction
- Two ongoing technological revolutions: genetic sequencing and genetic predictions.
- Hopes
- Precision medicine
- Preventive treatments
- And fears
- GATTACA
- Insurers cream-skimming (Genetic Information Nondiscrimination Act of 2008)
- Consumer gaming (adverse selection)
This paper: will selection be a problem?
Contributions
- Develop a methodology to quantify adverse selection from current genetic-prediction technology using genetic datasets.
- Extend the methodology to future prediction technology using heritability estimates and empirical regularities.
- Apply it to:
- Catastrophic illness insurance—a $100 billion industry
- Genetic and healthcare-utilization data from UK Biobank: 500,000 genomes plus NHS data
Example: prostate cancer (covariates)
Example: prostate cancer (covariates + current PGI)
Example: prostate cancer (genetics + true PGI)
Roadmap
- Background on genetics and catastrophic illness insurance
- Theory
- Case study: prostate cancer
- Wholesale application
Roadmap
- Background on genetics and catastrophic illness insurance
- Theory
- Case study: prostate cancer
- Wholesale application
The human DNA
- 3 billion base pairs across 23 chromosome pairs
- Most pairs are identical across individuals
- SNP: a location with differences in more than 1% of people
- About 600 million SNPs
- For most people, roughly 5 million SNPs differ from the reference
- Typical dataset: roughly 1 million SNPs
The human DNA
- 3 billion base pairs across 23 chromosome pairs
- Most pairs are identical across individuals
- SNP: a location with differences in more than 1% of people
- About 600 million SNPs
- For most people, roughly 5 million SNPs differ from the reference
- Typical dataset: roughly 1 million SNPs
The sequencing revolution
![Human Genome Project announcement]()
2003: Human Genome Project completed for approximately $3 billion.
![Consumer genetic sequencing advertisements]()
2023: Genome sequencing costs approximately $1,000; consumer genotyping costs as little as $79.
The sequencing revolution
The prediction revolution
![Galton's 1886 illustration of regression toward mediocrity in hereditary stature]()
Galton (1886), “Regression toward mediocrity in hereditary stature”
The prediction revolution
| Monogenic disorders (e.g., Huntington’s disease) |
≈100% |
| Height |
≈90% |
| Body Mass Index |
≈70% |
| Personality |
≈50% |
| Political behavior and attitudes |
≈40–60% |
| Educational attainment |
≈40% |
| Self-employment |
≈40% |
| Subjective well-being |
≈35% |
| Personal income |
≈30–40% |
| Decision-making anomalies |
≈30% |
| Portfolio choice |
≈30% |
| Economic preferences (risk, social preferences) |
≈20–30% |
Sources include Alford et al. (2005), Ashenfelter and Krueger (1994), Behrman and Taubman (1976), Cesarini et al. (2008–2011), and others.
Prediction and catastrophic insurance
How predictions work
- Most traits are explained by many genes with tiny individual effects—for example, breast cancer.
- Code each SNP by its number of adenine alleles: \(x_j \in \{0,1,2\}\).
- Run a regression: \[Y = \boldsymbol{\beta}\mathbf{X} + \varepsilon\]
- Then calculate a polygenic index (PGI) for any individual \(i\): \[G_i = \widehat{\boldsymbol{\beta}}\mathbf{X}_i\]
- This is a data-hungry procedure that improves with larger datasets and consortia.
Most traits are polygenic
![Distribution of the Alzheimer's disease PGI]()
Alzheimer’s disease
![Distribution of the breast cancer PGI]()
Breast cancer
![Distribution of the coronary artery disease PGI]()
Coronary artery disease
![Distribution of the colorectal cancer PGI]()
Colorectal cancer
![Distribution of the prostate cancer PGI]()
Prostate cancer
![Distribution of the schizophrenia PGI]()
Schizophrenia
![Distribution of the type 2 diabetes PGI]()
Type 2 diabetes
Standardized PGIs for the seven diseases studied in UK Biobank.
More data, better predictions (2011)
![Manhattan plot based on 9,394 cases and 12,462 controls]()
Ripke et al. (2011): 5 genome-wide significant sites.
More data, better predictions (2014)
![Manhattan plot based on 36,989 cases and 113,075 controls]()
Ripke et al. (2014): 108 genome-wide significant sites.
The polygenic index scaling law
- Prediction \(R^2\) scales predictably with GWAS sample size.
Dudbridge (2013); O’Connor et al. (2019); Harden & Koellinger (2020)
Predictions in practice
Predictions in practice
Adapted from MyHeritage.nl under fair use for research and education.
Institutional background: critical illness insurance (CII)
The contract
- One-time payout upon any covered condition
- \(\geq 80\%\) of claims from common cancers and (cardio-)vascular disease
- Medical underwriting sets renewable term premiums until end age (often 65)
The market
- \(\sim 100\) million policyholders worldwide; top-selling in Asia
- US: 6.3M+ policies (2022), outnumbering long-term care; among the fastest-growing insurance types
- In demand even with universal healthcare (UK, Canada, Australia)
Gatzert and Maegebier (2015); Gen Re (2023); Swiss Re Institute (2024)
Roadmap
- Background on genetics and catastrophic illness insurance
- Theory
- Case study: prostate cancer
- Wholesale application
Economic model
- A population of consumers whom insurers treat similarly may suffer a binary loss—for example, prostate cancer within a narrow risk class.
- Each consumer is characterized by four random variables: \[
(L, G_c, G_f, W).
\]
- \(L \in \{0,1\}\) is the binary loss.
- \(G_c\) is the currently available polygenic score.
- \(G_f\) is the perfect genetic predictor (the future polygenic score).
- \(W\) is non-genetic private information.
- The population is described by distribution \(\mathbb{P}\).
Measuring selection
Privately known risk
- Non-genetic risk: \[P = \mathbb{E}[L \mid W].\]
- Genetic risk: \[R = \mathbb{E}[L \mid G, W].\]
We expect little dispersion in non-genetic risk, more selection using the current PGI, and still more using the future PGI.
Selection depends on the entire risk distribution. For simplicity, we often report the implicit tax at a percentile (Hendren, 2013).
Genetic model
Key assumptions
- Probit. Loss follows a linear probit model in \(G_f\) and \(W\).
- Gaussian predictor. Conditional on \(W\), the true genetic predictor is Gaussian, with constant variance and a mean linear in \(W\).
- Noisy PGI. The current PGI is the true genetic predictor plus independent Gaussian noise: \[G_c = G_f + \varepsilon.\]
- Future predictive power. The predictive power of \(G_f\) is given by the known heritability estimate: \[R^2_{\mathrm{heritability}}.\]
Main theorem
Definition. Given the observed joint distribution
\[
\mathbb{P}_{\mathrm{data}}(L,G_c,W),
\]
the model is identified if there is a unique joint distribution
\[
\mathbb{P}(L,G_c,G_f,W)
\]
consistent with the data.
Theorem. Under assumptions 1–4 and additional technical conditions, the model is identified.
Proof idea
The model is
\[
\begin{aligned}
W &\sim \mathbb{P}_W, \\
G_f &= W\theta + V, \\
L &= \mathbf{1}\!\left\{G_f\beta_g + W\beta_w + \eta > 0\right\}.
\end{aligned}
\]
But we observe only a noisy current PGI:
\[
G_c = G_f + \varepsilon.
\]
Identifying parameters
- The variances match the current \(R^2\) and the future \(R^2\) implied by heritability estimates.
- \(\theta\) is straightforward: regress the current PGI \(G_c\) on \(W\).
- The remaining non-obvious parameters are the \(\beta\)s.
Key step. Conditional on \(W=w\) and \(G_c=g_c\), we observe a noisy signal of \(G_f\). By the Gaussian updating formula,
\[
G_f \mid (W=w,G_c=g_c)
\sim \mathrm{N}\!\left(a w\theta + b g_c,\,c\right).
\]
Roadmap
- Background on genetics and catastrophic illness insurance
- Theory
- Case study: prostate cancer
- Wholesale application
Data: UK Biobank
- A long-running genetic panel created for genetic epidemiology
- Approximately 500,000 volunteers plus extensions; our largest analyses have \(N \approx 400{,}000\)
- DNA data, questionnaires, exams, and NHS inpatient records
- About 850,000 SNPs per person
- Each row therefore has 850,000 variables \(x_j \in \{0,1,2\}\)
Taking the model to the data
- Consider prostate-cancer critical-illness insurance.
- A typical enrollee is around 35, and companies must renew at the same rate.
- If consumers always renew, the relevant loss is \(L=1\) if the consumer ever experiences the loss.
- Assume we can observe all 35-year-olds until death.
Non-genetic risk: \(\Pr(L=1)=f(W)\)
Current genetic risk: \(\Pr(L=1)=f(G,W)\)
Taking the model to the data
- Estimation with the current PGI is straightforward.
- Under our assumptions, the identification theorem lets us estimate the full model and evaluate selection using probit regressions.
- Additional engineering is required:
- Linear models
- Censored data
- Using data from people other than 35-year-olds
- Feature selection
Non-genetic risk
Current genetic risk
Future genetic risk: lower bound
Future genetic risk: twin upper bound
Non-genetic risk
Current genetic risk
Future genetic risk: lower bound
Future genetic risk: twin upper bound
Roadmap
- Background on genetics and catastrophic illness insurance
- Theory
- Case study: prostate cancer
- Wholesale application
Implicit taxes
Main results
The results are broadly consistent with the prostate-cancer case study.
- With current prediction technology and widespread testing, selection would be noticeable.
- With future prediction technology, selection would be crippling for many contracts.
- There is substantial variation across single-disease critical-illness insurance contracts.
Discussion
- Bundled contracts
- Incomplete take-up
- Equilibrium models
- Inverse selection—better-informed firms
- Changes in behavior
- Policy
Conclusion
Practical question
- Genetic predictions are likely to cause substantial adverse selection in some future insurance markets—but not for every risk.
Methodological contribution
- Methods to evaluate this question in different settings.
Future research
- Epidemiological applications
- Larger insurance markets: health and long-term care
- Capturing the benefits of this amazing technology while managing unintended consequences
Thank you!