Genetic Prediction and Adverse Selection

with Richard Karlsson Linnér and Jonathan Beauchamp

Eduardo Azevedo

Wharton

Introduction

  • Two ongoing technological revolutions: genetic sequencing and genetic predictions.
  • Hopes
    • Precision medicine
    • Preventive treatments
  • And fears
    • GATTACA
    • Insurers cream-skimming (Genetic Information Nondiscrimination Act of 2008)
    • Consumer gaming (adverse selection)

This paper: will selection be a problem?

Contributions

  1. Develop a methodology to quantify adverse selection from current genetic-prediction technology using genetic datasets.
  2. Extend the methodology to future prediction technology using heritability estimates and empirical regularities.
  3. Apply it to:
    • Catastrophic illness insurance—a $100 billion industry
    • Genetic and healthcare-utilization data from UK Biobank: 500,000 genomes plus NHS data

Example: prostate cancer (covariates)

Distribution of prostate-cancer risk using covariates

Example: prostate cancer (covariates + current PGI)

Distribution of prostate-cancer risk using covariates and the current PGI

Example: prostate cancer (genetics + true PGI)

Distribution of prostate-cancer risk using the future PGI at the twin-heritability bound

Roadmap

  1. Background on genetics and catastrophic illness insurance
  2. Theory
    • Economics
    • Genetics
  3. Case study: prostate cancer
  4. Wholesale application

Roadmap

  1. Background on genetics and catastrophic illness insurance
  2. Theory
    • Economics
    • Genetics
  3. Case study: prostate cancer
  4. Wholesale application

The human DNA

  • 3 billion base pairs across 23 chromosome pairs
  • Most pairs are identical across individuals
  • SNP: a location with differences in more than 1% of people
  • About 600 million SNPs
  • For most people, roughly 5 million SNPs differ from the reference
  • Typical dataset: roughly 1 million SNPs

Diagram of DNA base pairs

The human DNA

  • 3 billion base pairs across 23 chromosome pairs
  • Most pairs are identical across individuals
  • SNP: a location with differences in more than 1% of people
  • About 600 million SNPs
  • For most people, roughly 5 million SNPs differ from the reference
  • Typical dataset: roughly 1 million SNPs

Binary representation of genetic information

The sequencing revolution

Human Genome Project announcement

2003: Human Genome Project completed for approximately $3 billion.

Consumer genetic sequencing advertisements

2023: Genome sequencing costs approximately $1,000; consumer genotyping costs as little as $79.

The sequencing revolution

Decline in the cost per raw megabase of DNA sequence

The prediction revolution

Galton's 1886 illustration of regression toward mediocrity in hereditary stature
Galton (1886), “Regression toward mediocrity in hereditary stature”

The prediction revolution

Trait Heritability
Monogenic disorders (e.g., Huntington’s disease) ≈100%
Height ≈90%
Body Mass Index ≈70%
Personality ≈50%
Political behavior and attitudes ≈40–60%
Educational attainment ≈40%
Self-employment ≈40%
Subjective well-being ≈35%
Personal income ≈30–40%
Decision-making anomalies ≈30%
Portfolio choice ≈30%
Economic preferences (risk, social preferences) ≈20–30%
Sources include Alford et al. (2005), Ashenfelter and Krueger (1994), Behrman and Taubman (1976), Cesarini et al. (2008–2011), and others.

Prediction and catastrophic insurance

Current PGI R-squared and SNP and twin heritability estimates across catastrophic illnesses on a scale from zero to one hundred percent

How predictions work

  • Most traits are explained by many genes with tiny individual effects—for example, breast cancer.
  • Code each SNP by its number of adenine alleles: \(x_j \in \{0,1,2\}\).
  • Run a regression: \[Y = \boldsymbol{\beta}\mathbf{X} + \varepsilon\]
  • Then calculate a polygenic index (PGI) for any individual \(i\): \[G_i = \widehat{\boldsymbol{\beta}}\mathbf{X}_i\]
  • This is a data-hungry procedure that improves with larger datasets and consortia.

Most traits are polygenic

Distribution of the Alzheimer's disease PGI

Alzheimer’s disease

Distribution of the breast cancer PGI

Breast cancer

Distribution of the coronary artery disease PGI

Coronary artery disease

Distribution of the colorectal cancer PGI

Colorectal cancer

Distribution of the prostate cancer PGI

Prostate cancer

Distribution of the schizophrenia PGI

Schizophrenia

Distribution of the type 2 diabetes PGI

Type 2 diabetes

Standardized PGIs for the seven diseases studied in UK Biobank.

More data, better predictions (2011)

Manhattan plot based on 9,394 cases and 12,462 controls
Ripke et al. (2011): 5 genome-wide significant sites.

More data, better predictions (2014)

Manhattan plot based on 36,989 cases and 113,075 controls
Ripke et al. (2014): 108 genome-wide significant sites.

The polygenic index scaling law

  • Prediction \(R^2\) scales predictably with GWAS sample size.

Observed versus predicted polygenic-index accuracy across traits as GWAS sample size increases

Scaling curves relating GWAS discovery sample size to the predictive accuracy of polygenic scores

Dudbridge (2013); O’Connor et al. (2019); Harden & Koellinger (2020)

Predictions in practice

New York Times article about genetic testing and preventive treatment

Predictions in practice

Consumer genetic-testing interface

Consumer genetic risk score for type 2 diabetes

Adapted from MyHeritage.nl under fair use for research and education.

Predictions in practice

Undark logo

Article titled From a Fledgling Genetic Science, A Murky Market for Prediction

Genetics in the news

Social-media discussion of genetic risk

Article about genetic variants and financial backing

Article about consumer DNA databases

Institutional background: critical illness insurance (CII)

The contract

  • One-time payout upon any covered condition
  • \(\geq 80\%\) of claims from common cancers and (cardio-)vascular disease
  • Medical underwriting sets renewable term premiums until end age (often 65)

The market

  • \(\sim 100\) million policyholders worldwide; top-selling in Asia
  • US: 6.3M+ policies (2022), outnumbering long-term care; among the fastest-growing insurance types
  • In demand even with universal healthcare (UK, Canada, Australia)
Gatzert and Maegebier (2015); Gen Re (2023); Swiss Re Institute (2024)

Roadmap

  1. Background on genetics and catastrophic illness insurance
  2. Theory
    • Economics
    • Genetics
  3. Case study: prostate cancer
  4. Wholesale application

Economic model

  • A population of consumers whom insurers treat similarly may suffer a binary loss—for example, prostate cancer within a narrow risk class.
  • Each consumer is characterized by four random variables: \[ (L, G_c, G_f, W). \]
    • \(L \in \{0,1\}\) is the binary loss.
    • \(G_c\) is the currently available polygenic score.
    • \(G_f\) is the perfect genetic predictor (the future polygenic score).
    • \(W\) is non-genetic private information.
  • The population is described by distribution \(\mathbb{P}\).

Measuring selection

Privately known risk

  • Non-genetic risk: \[P = \mathbb{E}[L \mid W].\]
  • Genetic risk: \[R = \mathbb{E}[L \mid G, W].\]

We expect little dispersion in non-genetic risk, more selection using the current PGI, and still more using the future PGI.

Selection depends on the entire risk distribution. For simplicity, we often report the implicit tax at a percentile (Hendren, 2013).

Genetic model

Key assumptions

  1. Probit. Loss follows a linear probit model in \(G_f\) and \(W\).
  2. Gaussian predictor. Conditional on \(W\), the true genetic predictor is Gaussian, with constant variance and a mean linear in \(W\).
  3. Noisy PGI. The current PGI is the true genetic predictor plus independent Gaussian noise: \[G_c = G_f + \varepsilon.\]
  4. Future predictive power. The predictive power of \(G_f\) is given by the known heritability estimate: \[R^2_{\mathrm{heritability}}.\]

Main theorem

Definition. Given the observed joint distribution

\[ \mathbb{P}_{\mathrm{data}}(L,G_c,W), \]

the model is identified if there is a unique joint distribution

\[ \mathbb{P}(L,G_c,G_f,W) \]

consistent with the data.

Theorem. Under assumptions 1–4 and additional technical conditions, the model is identified.

Proof idea

The model is

\[ \begin{aligned} W &\sim \mathbb{P}_W, \\ G_f &= W\theta + V, \\ L &= \mathbf{1}\!\left\{G_f\beta_g + W\beta_w + \eta > 0\right\}. \end{aligned} \]

But we observe only a noisy current PGI:

\[ G_c = G_f + \varepsilon. \]

Identifying parameters

  • The variances match the current \(R^2\) and the future \(R^2\) implied by heritability estimates.
  • \(\theta\) is straightforward: regress the current PGI \(G_c\) on \(W\).
  • The remaining non-obvious parameters are the \(\beta\)s.

Key step. Conditional on \(W=w\) and \(G_c=g_c\), we observe a noisy signal of \(G_f\). By the Gaussian updating formula,

\[ G_f \mid (W=w,G_c=g_c) \sim \mathrm{N}\!\left(a w\theta + b g_c,\,c\right). \]

Reverse-shrinkage formula

Run a probit regression of loss on \(G_c\) and \(W\). Its coefficients \(\gamma_g\) and \(\gamma_w\) are related to the structural coefficients:

\[ \beta_g = \frac{\gamma_g}{\sqrt{a^2-\gamma_g^2c^2}}, \]

\[ \beta_w = \gamma_w\sqrt{1+c^2\beta_g^2}-b\theta\beta_g. \]

Key idea. Future genes that remain to be discovered are assumed to correlate with covariates in a similar way to currently discovered genes.

Roadmap

  1. Background on genetics and catastrophic illness insurance
  2. Theory
    • Economics
    • Genetics
  3. Case study: prostate cancer
  4. Wholesale application

Data: UK Biobank

  • A long-running genetic panel created for genetic epidemiology
  • Approximately 500,000 volunteers plus extensions; our largest analyses have \(N \approx 400{,}000\)
  • DNA data, questionnaires, exams, and NHS inpatient records
  • About 850,000 SNPs per person
  • Each row therefore has 850,000 variables \(x_j \in \{0,1,2\}\)

Taking the model to the data

  • Consider prostate-cancer critical-illness insurance.
  • A typical enrollee is around 35, and companies must renew at the same rate.
  • If consumers always renew, the relevant loss is \(L=1\) if the consumer ever experiences the loss.
  • Assume we can observe all 35-year-olds until death.

Non-genetic risk: \(\Pr(L=1)=f(W)\)
Current genetic risk: \(\Pr(L=1)=f(G,W)\)

Taking the model to the data

  • Estimation with the current PGI is straightforward.
  • Under our assumptions, the identification theorem lets us estimate the full model and evaluate selection using probit regressions.
  • Additional engineering is required:
    • Linear models
    • Censored data
    • Using data from people other than 35-year-olds
    • Feature selection

Non-genetic risk

Distribution of prostate-cancer risk using covariates

Current genetic risk

Distribution of prostate-cancer risk using covariates and the current PGI

Future genetic risk: lower bound

Distribution of prostate-cancer risk using the future PGI at the SNP-heritability bound

Future genetic risk: twin upper bound

Distribution of prostate-cancer risk using the future PGI at the twin-heritability bound

Non-genetic risk

Implicit tax by risk percentile using covariates

Current genetic risk

Implicit tax by risk percentile using the current PGI

Future genetic risk: lower bound

Implicit tax by risk percentile using the future PGI at the SNP-heritability bound

Future genetic risk: twin upper bound

Implicit tax by risk percentile using the future PGI at the twin-heritability bound

Comparison with community rating

Implicit tax using covariates under community rating across all risk classes

Roadmap

  1. Background on genetics and catastrophic illness insurance
  2. Theory
    • Economics
    • Genetics
  3. Case study: prostate cancer
  4. Wholesale application

Implicit taxes

Minimum implicit tax through the 80th risk percentile across contracts and prediction scenarios

Main results

The results are broadly consistent with the prostate-cancer case study.

  1. With current prediction technology and widespread testing, selection would be noticeable.
  2. With future prediction technology, selection would be crippling for many contracts.
  3. There is substantial variation across single-disease critical-illness insurance contracts.

Discussion

  • Bundled contracts
  • Incomplete take-up
  • Equilibrium models
  • Inverse selection—better-informed firms
  • Changes in behavior
  • Policy

Conclusion

Practical question

  • Genetic predictions are likely to cause substantial adverse selection in some future insurance markets—but not for every risk.

Methodological contribution

  • Methods to evaluate this question in different settings.

Future research

  • Epidemiological applications
  • Larger insurance markets: health and long-term care
  • Capturing the benefits of this amazing technology while managing unintended consequences
Thank you!