When biomarkers, assays, and protocols co-evolve

Cross-platform reproducibility, validation, and adaptive design — FDA CDRH OSEL Seminar

Naim Rashid, PhD
Associate Professor, Department of BiostatisticsUNC Gillings School of Global Public HealthMember, UNC Lineberger Comprehensive Cancer Centernaim@unc.edu  ·  naimurashid.github.io

2026-05-20

§1 — ADAPT: the case study

Metastatic cancer: why progress stays incremental

JNCI 2023

Despite precision-oncology advances, metastatic cancer remains a leading cause of death — population-level survival gains stay modest.

Why aren’t we seeing bigger gains? Part of the answer is structural — and it is what ADAPT targets:

  • Trials are fixed, not adaptive — they cannot re-route as resistance emerges
  • Biomarkers are read once at baseline, not followed continuously
  • Discovery is retrospective — findings arrive after the trial closes

The ARPA-H ADAPT opportunity

March 2024: ARPA-H launches ADAPT to close those three gaps.

The ARPA-H ask: Design an adaptive multi-line trial to discover, evaluate, and implement never-before-seen baseline and longitudinal biomarkers in real time — including N-of-1 arms (single-patient adaptive arms) — on a one-month submission timeline.

Three verbs — discover, evaluate, implement — that this talk will return to as a single arc.

Longitudinal and multi-modal biomarkers

Longitudinal: measurement before, during, after treatment; arms re-route as resistance emerges.

Multi-modal: single-gene → multi-platform predictive biomarkers.

The regulatory problem: biomarker discovery, assay transfer, validation, and protocol decision rules usually live in separate pipelines. ADAPT forces them into one system — and that coupling is the subject of this talk.

ADAPT’s structure — three Technical Areas, one of them EVOLVE

TA1 (biomarker engine) ↔︎ TA2 (the trials) ↔︎ TA3 (data infrastructure). TA2 sends samples to TA1; TA1 proposes candidate biomarkers back as new targeting arms. The TA numbers label the areas — they are not the flow order.

EVOLVE-BDT (NCT07340541): a biomarker-stratified multi-arm trial in 2L metastatic breast cancer — ~700 patients, 10–15 arms; TA1 adds arms as resistance targets emerge. Structural point: the arm flow; drugs and subgroups are illustrative.

§2 — The coupled system

The same specimen, a different call

A locked biomarker rule classifies a patient’s tumor sample. Call: biomarker-positive.

That output determines:

  • subtrial / monitoring path
  • interim monitoring statistic
  • efficacy / futility boundary crossing

Re-tested on the deployment platform, the same specimen gets a different call.

Thesis

When a biomarker call drives an adaptive trial’s decisions — patient routing, interim stopping — its uncertainty is no longer an assay property; it becomes part of the trial design, to be calibrated and evaluated with it.

Review-relevant object = the coupled system, not the biomarker in isolation.

The coupled review object

Layer Uncertainty Control
Biomarker rule threshold drift lock + held-out eval
Assay / platform cross-platform shift bridging study
Population prevalence shift scenario envelope
Monitoring weak endpoint link surrogacy evidence
Adaptation false promotion simulated OCs
Governance post-hoc changes pre-specification

Each layer carries an uncertainty. All propagate to operating characteristics.

The design must hold across the plausible range of these uncertainties — their envelope — not a single point estimate.

These six aspects bundle concerns that CDRH’s PCCP final guidance (December 2024) addresses separately. In ADAPT, they cannot be reviewed in isolation.

One arc: discovery → evaluation → calibration

The six uncertainty layers are not six separate problems. Each lives inside one of three movements — and each movement is also a device-evaluation question. This is the arc for the rest of the talk.

Arc movement The question it answers This talk Device-review counterpart
Discovery (§3) Does a biomarker survive the move to the trial’s measurement platform? Cross-platform reproducibility; rank-based transfer (RIF) Cross-deployment classifier drift
Evaluation (§4) What is the uncertainty in each call, and how does it propagate to trial decisions? Per-patient conformal uncertainty; PPoS; classification-error propagation Per-prediction uncertainty for AI/ML classifiers
Calibration (§5) Can we find a design whose operating characteristics hold across that uncertainty? Simulation-based constrained calibration (BATON) Premarket simulation for adaptive device trials

Discover, evaluate, implement — the ARPA-H verbs — read in two vocabularies at once.

§3 — Discovery: biomarkers across platforms

Three kinds of biomarker decision rule

Three kinds of biomarker enter ADAPT as decision rules — each mapping onto one of the figure’s readiness tiers:

  • a measured marker with an established cutpoint — pre-validated
  • a prediction-model score (PurIST, RIF, Oncotype) — needs calibration
  • a multimodal composite fused across platforms — needs development

The tier sets how much calibration each kind needs — but each, once it drives a decision, carries uncertainty into the trial design.

Our worked example below uses PurIST, a pancreatic-cancer subtype classifier — chosen because it has mature cross-platform validation data. The machinery is identical for EVOLVE’s breast-cancer biomarkers.

Cross-platform reproducibility: training versus deployment

The platform-transfer problem is distribution shift — a measurement/representation shift between the training and deployment platforms (the feature scale itself changes), not random noise. Locking the rule is the precondition for evaluation; PCCP then defines what controlled change is permitted after the locked snapshot.

Real ADAPT example. A classifier developed on RNA-seq, applied directly to Tempus XR:

  • Predicted values may be out of range / uninterpretable
  • Locked clinical thresholds are non-transferable
  • Classifier and cutpoints must be re-bridged before trial use

Bridging-study metrics: percent agreement and Cohen’s κ — the analytical-reproducibility layer.

Imaging / ML analogs: cross-scanner harmonization, cross-vendor MRI, model drift.

OSEL connection. Cross-platform bridging is the same review object as cross-deployment reproducibility for AI/ML devices under PCCP. In-trial validation does not substitute for it.

RIF: a rank-based answer to distribution shift

Rank Invariant Forest (RIF) — our lab’s rank-based classifier — classifies by gene-pair rank order, invariant to the monotone platform shifts that break absolute-expression models.

Each split is a gene-pair rank comparison (e.g., GATA6 > S100A2?). Ranks survive monotone platform shifts; absolute expression levels do not.

Worked example (non-ADAPT, pancreatic). Vs the CLIA-NS reference: PurIST holds up on its native RNA-seq (κ = 0.66) but degrades on Twist exome (κ = 0.60). On Twist exome, RIF reaches κ = 0.72 — closing the gap.

Rank-based transfer handles distribution shift — not the per-patient uncertainty that §4 takes up.

§4 — Evaluation: per-patient uncertainty

The uncertainty-quantification (UQ) gap for model-derived biomarkers

A device may output a calibrated score — but the decision protocol that consumes it collapses to a hard call. The gap is between the score and the action.

  • Most classifiers output a point probability — the decision is made on the dichotomized call (PurIST, for example, splits Basal-like from Classical at a 0.5 cutoff)
  • “Lean” / “borderline” zones are ad hoc cutoffs

The harm: at decision time, a borderline call and a confident call look identical. The clinician acts on a hard label with no quantified confidence.

For OSEL: AI/ML devices with a PCCP that lock thresholds without a per-prediction UQ layer give reviewers no visibility into which calls were borderline.

RIF response: a data-driven Mixed / no-call zone

Within PurIST-Classical (n = 660), split by RIF conformal call.

  • Conformal prediction wraps a score into a rule with a finite-sample coverage guarantee
  • RIF carves a data-driven Mixed zone — width set by the data to a 92.5% coverage target
  • Binary PurIST collapses all of these into “Classical”
RIF conformal stratum HR 95% CI
Confident Classical ref
Mixed 1.49 0.96–2.33
Confident Basal 2.30 1.62–3.25

Mixed survival sits between the confident classes — intermediate, not noise.

RIF adds two strata — each changes a different decision

Conformal RIF splits the binary “Classical” call into three strata — the two new ones drive two distinct actions.

Correction 1 — reclassification

PurIST calls all 660 patients “Classical”; RIF’s conformal call splits them:

PurIST call Conf. Classical Mixed Conf. Basal
Classical (n = 660) 577 (87%) 32 (5%) 51 (8%)

The 8% reclassified Confident BasalHR 2.30, 9-month OS gap — are wrong by default under binary PurIST.

Correction 2 — abstention (Mixed)

RIF flags an intermediate stratum, confidently neither class — HR 1.49 (CI crosses 1).

Action: orthogonal work-up; defer therapy until disambiguated.

Abstention is a design knob, not a free lever — it needs orthogonal testing available and reimbursed.

OSEL principle. A pre-specified abstention layer is a risk-control primitive — coverage-guaranteed uncertainty added without revisiting the locked classifier.

Evaluating a proposed biomarker: PPoS

Before a biomarker earns a targeted intervention (an N-of-1 or expanded arm), the trial must show prospectively that B+ and B− differ in survival (progression-free survival, PFS). PPoS — predictive probability of success — is, at each interim, the probability the final analysis would declare a B+/B− difference if the trial ran to its planned end (from the current posterior).

Detection — a GO when B+/B− truly differ; false promotion — a GO when they don’t. (Figure: “False Positive” = false promotion; “False Negative” = missed biomarker; thresholds set by BATON, §5.) Next: how classification error moves both.

Biomarker uncertainty propagates to screening power

The biomarker call itself is not error-free. Aggregated over an arm, classification error degrades the detection and false-promotion probabilities just defined.

Nondifferential classification error (figure: three endpoints — PFS, ORR, ctDNA): false-promotion barely moves (Δ < 3 pp), but detection drops sharply — on ORR, −9 pp at CE = 10%, −13 pp at CE = 15%. Misclassification attenuates the true effect and cuts power.

The §2 envelope, made concrete: CE → detection probability — a real loss of screening power, not an inflation of false promotion.

When propagation becomes bias: differential drift

The dilution argument holds only if classification error is nondifferential. When error correlates with site, outcome, treatment, or platform, the argument breaks — and bias returns.

Differential failure modes:

  • Vendor-by-center differences
  • Batch effects correlated with outcome timing
  • Prevalence shift across sites
  • Sample-QC concentrated in one stratum

The transferable diagnostic — for OSEL: Is observed drift nondifferential with respect to subpopulation, outcome, site, or deployment environment?

  • If yes → loss of sensitivity dominates
  • If no → bias and false-positive control return as central review concerns

This is the post-market drift question for AI/ML devices with a PCCP — same diagnostic, different submission type.

§5 — Calibration: trial design with BATON

Simulation-based calibration: the operating standard

Regulatory posture. For Bayesian / adaptive designs with no closed-form operating characteristics, FDA-facing justification of thresholds and stopping rules is established through clinical trial simulation — per FDA’s Bayesian and adaptive-design guidance.

Stays human: endpoint choice, prior specification, benefit-risk judgment.

The framework extends. Because calibration is simulation-based, biomarker uncertainty can enter the loop directly — uncertain routing and outcome association are simulable like any design element. Bringing them inside the same search is the natural extension (§6).

BATON — our simulation-based design-calibration method — makes the threshold-search reproducible and auditable. Judgment stays with the committee; the search becomes a documented procedure with a stress-test envelope.

BATON: smart search through design space

  • BATON learns a surrogate map of the design space, then samples where feasible designs are likely
  • Reaches a comparable optimum in ~150 evaluations — far fewer than grid or the 2,000-design random search (6.8% feasible)
  • The same Bayesian optimization used for expensive black-box tuning
  • Same review object as simulation-calibrated change-control rules for AI/ML devices

Panels above: a schematic 2-D illustration of the search contrast — not the EVOLVE design counts.

In development — BATON-C extends the search to calibrate robustly across an envelope of uncertain trial assumptions (accrual, hazards, priors), not a single fixed point.

Validation — unit test: matches Simon optimal

12 Simon two-stage designs — a standard phase-II benchmark with known optima; no Simon-specific knowledge given to BATON.

  • 12 / 12 satisfy power ≥ 0.80 AND type I ≤ 0.10
  • \(E_0[N]\) is within ±5% of the Simon optimum in most benchmarks; where it differs, Simon is modestly more efficient — BATON never beats the known optimum
  • BATON recovers the known Simon optima without being told which problems are closed-form

The integration test — EVOLVE-like calibration with no closed-form ground truth — is §3–§5 of this talk.

Validation — reproducibility across independent runs

10 independent runs of the same calibration problem.

  • Coefficient of variation of \(E_0[N]\) < 1% across runs
  • Convergence in ~150 iterations
  • BOBYQA (a standard derivative-free optimizer): 30 runs → zero feasible

Calibration exposes trade-offs before launch

An illustrative single-arm scenario: three of the five philosophies located in \((N_{\max}, E_0[N])\) space.

Same constraints, different objectives:

Philosophy Objective
Optimal min \(E_0[N]\)
Minimax min \(N_{\max}\)
Fleming min avg \(E[N]\)
Alt-Optimal min \(E_1[N]\)
Admissible Pareto frontier

The point isn’t a “best” design — it’s an auditable trade-off space the steering committee can interrogate before launch.

Operational impact — a defensible calibration record

EVOLVE calibration comparison:

Manual BATON
Audit trail Incomplete Full
Stress-test envelope Ad-hoc Pre-specified
Reproducibility Person-dependent Reproducible (fixed code + seeds)
Time 2 weeks 4 hours
Power achieved 0.78 0.83

The audit trail is the propagation record that closes back to the §2 envelope.

Downstream use:

  • Pareto frontier informs steering-committee decisions
  • Audit trail supports FDA documentation
  • Interim timing structures the Data Safety Monitoring Board (DSMB) schedule

§6 — Translation to OSEL

The synthesis: one arc, two vocabularies

The talk traced one arc — discovery → evaluation → calibration — the same three ARPA-H verbs. Read in two vocabularies, it is the same arc for clinical trials and for device review.

Arc movement This talk’s framing The device-review counterpart
Discovery Cross-platform reproducibility under distribution shift Cross-deployment classifier drift
Evaluation Per-patient uncertainty (conformal) + how it propagates Per-prediction uncertainty for AI/ML classifiers
Calibration BATON: simulation-based design calibration Premarket simulation for adaptive device trials

EVOLVE (NCT07340541, ARPA-H ADAPT) is the live deployment.

Open questions and collaboration invitation

Open directions, each live for OSEL’s AI program:

  1. Robustness by construction. RIF survives platform shift via rank features. Can that carry to imaging and multi-modal devices — and how is robustness evaluated, not just detected post-market?
  2. Simulation-based calibration as a regulatory tool. BATON calibrates operating characteristics by black-box simulation. What credibility standard lets a reviewer rely on it for device change-control plans?
  3. Conformal guarantees under deployment shift. Per-prediction coverage assumes exchangeability — which shift breaks. What should conformal guidance for AI/ML classifiers look like?
  4. Joint biomarker-and-design calibration. The abstention level \(\alpha\) is admissible to the same BATON search as the trial parameters. What would an end-to-end joint calibration — and the evidence a reviewer needs to rely on it — require?

Open to discussing or collaborating on any of these.

Acknowledgments + Questions

Rashid Lab (UNC)

  • Amber Young — PhD student; led BATON design & validation

ADAPT Biomarker UQ Working Group

  • Stephen Bates (MIT)
  • Ying Yuan (MD Anderson)
  • Qian Cao (FDA CDRH)

EVOLVE Statistical Working Group

  • Susan Hilsenbeck (Baylor College of Medicine)
  • Nabihah Tayoub (DFCI)
  • Ruizhe Chen (Johns Hopkins)

EVOLVE multiple-PI (MPI) team

  • Lisa Carey (UNC)
  • Eric Winer (Yale)
  • Antonio Wolff (Johns Hopkins)
  • Ian Krop (Yale)
  • Charles Perou (UNC)

Funding — ARPA-H 140D042590009 · DOD HT9425241103100121 · NCI P50-CA058223 · NCI P50-CA257911 · NCI U01-CA274298

Contactnaim@unc.edu · naimurashid.github.io · X · Bluesky · GitHub · Google Scholar · LinkedIn

Backup slides

B1 · Mechanism — why nondifferential CE is asymmetric (dilution without bias)

Misclassification is approximately symmetric across B+/B−.

  • Under null: swap is invisible → false-promotion probability unchanged.
  • Under alternative: swapping dilutes the effect without biasing the test statistic → power down, Type I unchanged.

Berkson-style assignment-error (Pepe 2001).

Quantitatively:

  • 2×2 confusion matrix (per-stratum CE)
  • Dilution: \(\Delta_{obs} = (1 - 2 \cdot CE) \cdot \Delta_{true}\)
  • Power loss curves at CE = 10/15/20%

B2 · Clinical payoff — what binary calls miss

Within PurIST = Classical patients, RIF-flagged Basal-like calls — the OS-evaluable cross-cohort replication of the §4 reclassification result — show:

  • HR 2.21 (95% CI 1.58–3.11, p = 4 × 10⁻⁶) — OS-evaluable subset, 54 of 660 across 12 cohorts
  • Median OS gap 9 months (15 vs 24)
  • ~8% of OS-evaluable patients (54 of 660)
  • Mixed group: HR ≈ 1.49 — clinically intermediate survival, between the two confident strata

Note the two HR figures are different objects, not a discrepancy: HR 2.30 (CI 1.62–3.25) on the §4 “RIF response” slide is the Confident Basal stratum of the single n = 660 cohort (n = 51 on that slide’s survival curve); HR 2.21 here is the OS-evaluable subset replicated across 12 cohorts (54 patients). Both are correct; cite whichever matches the question.

Available if the conformal-abstention slide draws questions about clinical impact.

B3 · Platform-agnostic translation — broader cross-domain table

The framework is platform-agnostic. Same problems, same machinery, different domains.

Genomics example (this talk) Imaging / ML analog (this audience)
PurIST cross-platform CE Cross-scanner ML classifier drift
Tempus vs Guardant vs FoundationOne Siemens vs GE vs Philips
Cross-assay ctDNA harmonization Cross-vendor MRI harmonization
Split-sample bridging study Phantom / standards-based calibration
Conformal Mixed / no-call abstention Continual-learning device monitoring under PCCP
Locked vs recalibratable classifier Locked vs PCCP-modified algorithm

B4 · Random-search baseline for EVOLVE calibration

For perspective on why BATON is needed:

  • Random search of 2,000 designs → only 6.8% feasible
  • Grid search in 6–13 dims → zero designs with power ≥ 0.80
  • Manual calibration of EVOLVE → 2 weeks, power 0.78 (failed)
  • BATON → ~150 designs, hours, power 0.83 with audit trail

The feasibility region is narrow and non-convex; brute-force methods fail. (Random search and the grid covered different candidate sets and resolutions; the shared finding is that the feasible region is easily missed by unguided search.)

B5 · FDA Bayesian / adaptive guidance — context for Q&A

FDA guidance on Bayesian and adaptive designs has long treated clinical-trial simulation as the established way to characterize operating characteristics — type I error, power, expected sample size, stopping probabilities — when a design has no tractable closed-form properties.

Endpoint choice, prior specification, and benefit-risk judgment remain human decisions: simulation calibrates thresholds, it does not replace that judgment.

B6 · ARPA-H ADAPT program context

For context on the scale and structure of the program:

  • 5 modeling groups (TA1) — biomarker discovery / development
  • 2 platform performers (TA3) — data infrastructure
  • 3 disease platforms (TA2):
    • Lung
    • Colon
    • Breast (EVOLVE)
  • ~1-month submission timeline at program launch
  • N-of-1 capacity embedded in the trial design