Cross-platform reproducibility, validation, and adaptive design — FDA CDRH OSEL Seminar
2026-05-20
JNCI 2023
Despite precision-oncology advances, metastatic cancer remains a leading cause of death — population-level survival gains stay modest.
Why aren’t we seeing bigger gains? Part of the answer is structural — and it is what ADAPT targets:
March 2024: ARPA-H launches ADAPT to close those three gaps.
The ARPA-H ask: Design an adaptive multi-line trial to discover, evaluate, and implement never-before-seen baseline and longitudinal biomarkers in real time — including N-of-1 arms (single-patient adaptive arms) — on a one-month submission timeline.
Three verbs — discover, evaluate, implement — that this talk will return to as a single arc.
Longitudinal: measurement before, during, after treatment; arms re-route as resistance emerges.
Multi-modal: single-gene → multi-platform predictive biomarkers.
The regulatory problem: biomarker discovery, assay transfer, validation, and protocol decision rules usually live in separate pipelines. ADAPT forces them into one system — and that coupling is the subject of this talk.
TA1 (biomarker engine) ↔︎ TA2 (the trials) ↔︎ TA3 (data infrastructure). TA2 sends samples to TA1; TA1 proposes candidate biomarkers back as new targeting arms. The TA numbers label the areas — they are not the flow order.
EVOLVE-BDT (NCT07340541): a biomarker-stratified multi-arm trial in 2L metastatic breast cancer — ~700 patients, 10–15 arms; TA1 adds arms as resistance targets emerge. Structural point: the arm flow; drugs and subgroups are illustrative.
A locked biomarker rule classifies a patient’s tumor sample. Call: biomarker-positive.
That output determines:
Re-tested on the deployment platform, the same specimen gets a different call.
Thesis
When a biomarker call drives an adaptive trial’s decisions — patient routing, interim stopping — its uncertainty is no longer an assay property; it becomes part of the trial design, to be calibrated and evaluated with it.
Review-relevant object = the coupled system, not the biomarker in isolation.
| Layer | Uncertainty | Control |
|---|---|---|
| Biomarker rule | threshold drift | lock + held-out eval |
| Assay / platform | cross-platform shift | bridging study |
| Population | prevalence shift | scenario envelope |
| Monitoring | weak endpoint link | surrogacy evidence |
| Adaptation | false promotion | simulated OCs |
| Governance | post-hoc changes | pre-specification |
Each layer carries an uncertainty. All propagate to operating characteristics.
The design must hold across the plausible range of these uncertainties — their envelope — not a single point estimate.
These six aspects bundle concerns that CDRH’s PCCP final guidance (December 2024) addresses separately. In ADAPT, they cannot be reviewed in isolation.
The six uncertainty layers are not six separate problems. Each lives inside one of three movements — and each movement is also a device-evaluation question. This is the arc for the rest of the talk.
| Arc movement | The question it answers | This talk | Device-review counterpart |
|---|---|---|---|
| Discovery (§3) | Does a biomarker survive the move to the trial’s measurement platform? | Cross-platform reproducibility; rank-based transfer (RIF) | Cross-deployment classifier drift |
| Evaluation (§4) | What is the uncertainty in each call, and how does it propagate to trial decisions? | Per-patient conformal uncertainty; PPoS; classification-error propagation | Per-prediction uncertainty for AI/ML classifiers |
| Calibration (§5) | Can we find a design whose operating characteristics hold across that uncertainty? | Simulation-based constrained calibration (BATON) | Premarket simulation for adaptive device trials |
Discover, evaluate, implement — the ARPA-H verbs — read in two vocabularies at once.
Three kinds of biomarker enter ADAPT as decision rules — each mapping onto one of the figure’s readiness tiers:
The tier sets how much calibration each kind needs — but each, once it drives a decision, carries uncertainty into the trial design.
Our worked example below uses PurIST, a pancreatic-cancer subtype classifier — chosen because it has mature cross-platform validation data. The machinery is identical for EVOLVE’s breast-cancer biomarkers.
The platform-transfer problem is distribution shift — a measurement/representation shift between the training and deployment platforms (the feature scale itself changes), not random noise. Locking the rule is the precondition for evaluation; PCCP then defines what controlled change is permitted after the locked snapshot.
Real ADAPT example. A classifier developed on RNA-seq, applied directly to Tempus XR:
Bridging-study metrics: percent agreement and Cohen’s κ — the analytical-reproducibility layer.
Imaging / ML analogs: cross-scanner harmonization, cross-vendor MRI, model drift.
OSEL connection. Cross-platform bridging is the same review object as cross-deployment reproducibility for AI/ML devices under PCCP. In-trial validation does not substitute for it.
Rank Invariant Forest (RIF) — our lab’s rank-based classifier — classifies by gene-pair rank order, invariant to the monotone platform shifts that break absolute-expression models.
Each split is a gene-pair rank comparison (e.g., GATA6 > S100A2?). Ranks survive monotone platform shifts; absolute expression levels do not.
Worked example (non-ADAPT, pancreatic). Vs the CLIA-NS reference: PurIST holds up on its native RNA-seq (κ = 0.66) but degrades on Twist exome (κ = 0.60). On Twist exome, RIF reaches κ = 0.72 — closing the gap.
Rank-based transfer handles distribution shift — not the per-patient uncertainty that §4 takes up.
A device may output a calibrated score — but the decision protocol that consumes it collapses to a hard call. The gap is between the score and the action.
The harm: at decision time, a borderline call and a confident call look identical. The clinician acts on a hard label with no quantified confidence.
For OSEL: AI/ML devices with a PCCP that lock thresholds without a per-prediction UQ layer give reviewers no visibility into which calls were borderline.
Within PurIST-Classical (n = 660), split by RIF conformal call.
| RIF conformal stratum | HR | 95% CI |
|---|---|---|
| Confident Classical | ref | — |
| Mixed | 1.49 | 0.96–2.33 |
| Confident Basal | 2.30 | 1.62–3.25 |
Mixed survival sits between the confident classes — intermediate, not noise.
Conformal RIF splits the binary “Classical” call into three strata — the two new ones drive two distinct actions.
Correction 1 — reclassification
PurIST calls all 660 patients “Classical”; RIF’s conformal call splits them:
| PurIST call | Conf. Classical | Mixed | Conf. Basal |
|---|---|---|---|
| Classical (n = 660) | 577 (87%) | 32 (5%) | 51 (8%) |
The 8% reclassified Confident Basal — HR 2.30, 9-month OS gap — are wrong by default under binary PurIST.
Correction 2 — abstention (Mixed)
RIF flags an intermediate stratum, confidently neither class — HR 1.49 (CI crosses 1).
Action: orthogonal work-up; defer therapy until disambiguated.
Abstention is a design knob, not a free lever — it needs orthogonal testing available and reimbursed.
OSEL principle. A pre-specified abstention layer is a risk-control primitive — coverage-guaranteed uncertainty added without revisiting the locked classifier.
Before a biomarker earns a targeted intervention (an N-of-1 or expanded arm), the trial must show prospectively that B+ and B− differ in survival (progression-free survival, PFS). PPoS — predictive probability of success — is, at each interim, the probability the final analysis would declare a B+/B− difference if the trial ran to its planned end (from the current posterior).
Detection — a GO when B+/B− truly differ; false promotion — a GO when they don’t. (Figure: “False Positive” = false promotion; “False Negative” = missed biomarker; thresholds set by BATON, §5.) Next: how classification error moves both.
The biomarker call itself is not error-free. Aggregated over an arm, classification error degrades the detection and false-promotion probabilities just defined.
Nondifferential classification error (figure: three endpoints — PFS, ORR, ctDNA): false-promotion barely moves (Δ < 3 pp), but detection drops sharply — on ORR, −9 pp at CE = 10%, −13 pp at CE = 15%. Misclassification attenuates the true effect and cuts power.
The §2 envelope, made concrete: CE → detection probability — a real loss of screening power, not an inflation of false promotion.
The dilution argument holds only if classification error is nondifferential. When error correlates with site, outcome, treatment, or platform, the argument breaks — and bias returns.
Differential failure modes:
The transferable diagnostic — for OSEL: Is observed drift nondifferential with respect to subpopulation, outcome, site, or deployment environment?
This is the post-market drift question for AI/ML devices with a PCCP — same diagnostic, different submission type.
The open question is now calibration: can we find a design that survives all of it?
Bayesian adaptive trials with biomarker-driven adaptation have no closed-form operating characteristics — each design \(\boldsymbol{\theta}\) maps by simulation to \(\boldsymbol{\theta} \to (\text{Power}, \text{Type I}, E_0[N], E_1[N], N_{\max})\).
Goal: minimize expected enrollment under regulatory constraints — over 6–13 interacting parameters that cannot be tuned in isolation.
The manual loop — choose → simulate → check → adjust → repeat — does not scale: random search is 6.8% feasible (2,000 designs); grid search yields zero designs at power ≥ 0.80; manual EVOLVE calibration took 2 weeks and failed (power 0.78).
Three governance gaps remain: local search, person-dependent designs, no audit trail.
A guided, simulation-based search is needed — one that is reproducible and auditable, not just fast.
Regulatory posture. For Bayesian / adaptive designs with no closed-form operating characteristics, FDA-facing justification of thresholds and stopping rules is established through clinical trial simulation — per FDA’s Bayesian and adaptive-design guidance.
Stays human: endpoint choice, prior specification, benefit-risk judgment.
The framework extends. Because calibration is simulation-based, biomarker uncertainty can enter the loop directly — uncertain routing and outcome association are simulable like any design element. Bringing them inside the same search is the natural extension (§6).
BATON — our simulation-based design-calibration method — makes the threshold-search reproducible and auditable. Judgment stays with the committee; the search becomes a documented procedure with a stress-test envelope.
Panels above: a schematic 2-D illustration of the search contrast — not the EVOLVE design counts.
In development — BATON-C extends the search to calibrate robustly across an envelope of uncertain trial assumptions (accrual, hazards, priors), not a single fixed point.
12 Simon two-stage designs — a standard phase-II benchmark with known optima; no Simon-specific knowledge given to BATON.
The integration test — EVOLVE-like calibration with no closed-form ground truth — is §3–§5 of this talk.
10 independent runs of the same calibration problem.
An illustrative single-arm scenario: three of the five philosophies located in \((N_{\max}, E_0[N])\) space.
Same constraints, different objectives:
| Philosophy | Objective |
|---|---|
| Optimal | min \(E_0[N]\) |
| Minimax | min \(N_{\max}\) |
| Fleming | min avg \(E[N]\) |
| Alt-Optimal | min \(E_1[N]\) |
| Admissible | Pareto frontier |
The point isn’t a “best” design — it’s an auditable trade-off space the steering committee can interrogate before launch.
EVOLVE calibration comparison:
| Manual | BATON | |
|---|---|---|
| Audit trail | Incomplete | Full |
| Stress-test envelope | Ad-hoc | Pre-specified |
| Reproducibility | Person-dependent | Reproducible (fixed code + seeds) |
| Time | 2 weeks | 4 hours |
| Power achieved | 0.78 | 0.83 |
The audit trail is the propagation record that closes back to the §2 envelope.
Downstream use:
The talk traced one arc — discovery → evaluation → calibration — the same three ARPA-H verbs. Read in two vocabularies, it is the same arc for clinical trials and for device review.
| Arc movement | This talk’s framing | The device-review counterpart |
|---|---|---|
| Discovery | Cross-platform reproducibility under distribution shift | Cross-deployment classifier drift |
| Evaluation | Per-patient uncertainty (conformal) + how it propagates | Per-prediction uncertainty for AI/ML classifiers |
| Calibration | BATON: simulation-based design calibration | Premarket simulation for adaptive device trials |
EVOLVE (NCT07340541, ARPA-H ADAPT) is the live deployment.
Open directions, each live for OSEL’s AI program:
Open to discussing or collaborating on any of these.
Rashid Lab (UNC)
ADAPT Biomarker UQ Working Group
EVOLVE Statistical Working Group
EVOLVE multiple-PI (MPI) team
Funding — ARPA-H 140D042590009 · DOD HT9425241103100121 · NCI P50-CA058223 · NCI P50-CA257911 · NCI U01-CA274298
Contact — naim@unc.edu · naimurashid.github.io · X · Bluesky · GitHub · Google Scholar · LinkedIn
Misclassification is approximately symmetric across B+/B−.
Berkson-style assignment-error (Pepe 2001).
Quantitatively:
Within PurIST = Classical patients, RIF-flagged Basal-like calls — the OS-evaluable cross-cohort replication of the §4 reclassification result — show:
Note the two HR figures are different objects, not a discrepancy: HR 2.30 (CI 1.62–3.25) on the §4 “RIF response” slide is the Confident Basal stratum of the single n = 660 cohort (n = 51 on that slide’s survival curve); HR 2.21 here is the OS-evaluable subset replicated across 12 cohorts (54 patients). Both are correct; cite whichever matches the question.
Available if the conformal-abstention slide draws questions about clinical impact.
The framework is platform-agnostic. Same problems, same machinery, different domains.
| Genomics example (this talk) | Imaging / ML analog (this audience) |
|---|---|
| PurIST cross-platform CE | Cross-scanner ML classifier drift |
| Tempus vs Guardant vs FoundationOne | Siemens vs GE vs Philips |
| Cross-assay ctDNA harmonization | Cross-vendor MRI harmonization |
| Split-sample bridging study | Phantom / standards-based calibration |
| Conformal Mixed / no-call abstention | Continual-learning device monitoring under PCCP |
| Locked vs recalibratable classifier | Locked vs PCCP-modified algorithm |
For perspective on why BATON is needed:
The feasibility region is narrow and non-convex; brute-force methods fail. (Random search and the grid covered different candidate sets and resolutions; the shared finding is that the feasible region is easily missed by unguided search.)
FDA guidance on Bayesian and adaptive designs has long treated clinical-trial simulation as the established way to characterize operating characteristics — type I error, power, expected sample size, stopping probabilities — when a design has no tractable closed-form properties.
Endpoint choice, prior specification, and benefit-risk judgment remain human decisions: simulation calibrates thresholds, it does not replace that judgment.
For context on the scale and structure of the program: