Paper 1 | An evidence-quality spine

Validity Gates Expose Limits of Longitudinal Multimodal Parkinson Disease Prediction

A model does not earn credibility from one favorable score. We ask seven separate questions: Was the future hidden? Were people and sites separated? Is the gain large enough to matter? Does it survive calendar change? Can uncertainty identify safer forecasts? Does the result travel to another cohort?

Harsh Milind Tirhekar, Priyanshi Yadav, Chandrajit Bajaj

5,404 future-only cutoffs from 1,058 PPMI participants 8,000 controlled gate-stress experiments One small internal signal; no deployment claim

The result at a glance

Seven questions produce seven verdicts, not one headline score

Each gate protects a different claim. Passing chronology cannot rescue failed transport, and correct interval coverage cannot substitute for useful error ranking.

ChronologyPass
ParticipantPass
SitePass
MagnitudeInformation only
CalendarFail
Selective UQAbstain
PDBP transportFail

The scientific output is the decision boundary. Dated multimodal history contains a reproducible internal increment, but the effect is clinically tiny, temporally fragile, and not externally transportable in the available export. The protocol therefore stops the claim before deployment language begins.

Study design

Every forecast begins at a dated cutoff

Imagine pausing one patient's record on a particular day. Only measurements already available at that moment may describe the state. Their age and absence remain visible; the future motor examination, 9 to 15 months later, stays hidden until evaluation.

1. Observed historyExam, wearable, DaT-SPECT, biomarker, and dated trends
2. LandmarkThe state is frozen using only information already available
3. Held-out futureThe later motor examination is revealed only for evaluation
One patient record, separated into observed history and held-out future
Conceptual same-patient clinical history, multimodal landmark review, and future motor examination separated by a held-out boundary
Conceptual medical illustration, not a participant record. The visible separation expresses the temporal firewall enforced programmatically in the analysis below.
Future-only longitudinal design and independent validity gates
A future-only patient timeline followed by separate participant, site, calendar, uncertainty, and transport gates
The model receives each source's most recent eligible value, age, availability, and observation context. The timeline is a protocol schematic, not an individual record. Every gate permits only its own claim.

Why ordinary validation is insufficient

A repeated clinical history is not a flat table

The same person appears at multiple visits. A scan may be months older than the examination it accompanies. Missingness may encode a study protocol rather than biology. Medication state changes motor scores, and a new enrollment era can change both who is measured and how sources are collected.

These are different scientific risks. The benchmark therefore assigns each one a distinct population, uncertainty interval, and promotion rule.

Analysis Scale Question it answers
Rolling-origin benchmark 1,058 people; 5,404 cutoffs Do dated sources add internal information when every person stays in one fold?
Strict-calendar challenge 442 later-era participants Does the frozen relation survive a later enrollment and observation regime?
Site-disjoint challenge 951 people across 52 sites Is the increment dependent on acquisition center?
Gate stress test 8,000 controlled experiments Does the protocol suppress false promotion and detect leakage or shift?
PDBP clinical transport 77 calibration; 156 untouched test participants Does a frozen clinical model travel to an independent cohort?

Internal result

The signal is measurable; the improvement is not clinically persuasive

Performance under progressively stricter validation
Held-out performance across participant, site, enrollment-era, and strict-calendar evaluations
Dated DaT-SPECT and MoCA history improve internal prediction over clinical trajectory by delta R-squared 0.015 and delta MAE -0.060 MDS-UPDRS Part III points. The interval supports a small information increment, but it misses the prespecified promotion magnitude and becomes uncertain under strict calendar testing.
0.015Internal delta R-squared
-0.060Delta MAE in motor points
Information onlyPrespecified magnitude verdict

A statistically stable improvement can still be practically negligible. The 0.060 point MAE reduction does not establish individual prognosis, patient-management value, or clinical-trial enrichment.

Uncertainty and stress testing

An interval may cover well and still fail to identify safer predictions

Coverage asks whether the future falls inside an interval. Selective prediction asks a harder question: when the model says it is more certain, is it actually less wrong?

Selective risk, participant coverage, and gate behavior
Selective-risk curves, participant-level interval coverage, and controlled gate behavior
Source-aware 90% intervals covered every eligible forecast for 90.6% of participants, yet interval width did not rank future error better than the clinical reference and reversed under calendar shift. The locked selective-prediction verdict is ABSTAIN. Across 8,000 controlled experiments, the full gate limited null promotion to 0.35%, detected constructed leakage, and blocked 94.75% of calendar reversals.

External transport

A large archive is not the same as a compatible cohort

PDBP contains valuable Parkinson disease data, but a transport claim requires the same diagnosis, modalities, timing, endpoint, and eligible participants to intersect. File presence alone is not validation.

Model-capacity and PPMI-to-PDBP transport stress tests
Model-capacity and external-transport stress tests from PPMI to PDBP
A six-variable clinical model achieved internal PPMI R-squared 0.134 but fell to -0.077 under raw PDBP transport. Recalibration reached R-squared 0.005 and did not beat the PDBP mean reference on MAE. The local imaging export contained no eligible diagnosis-modality-time intersection, so this is a failed transport and interoperability result, not external multimodal validation.

Claim contract

What the evidence permits us to say

Gate Verdict Permitted statement
Future-only chronologyPassEvery predictor precedes its endpoint
Participant separationPassThe internal signal is not produced by repeated-person leakage
Site separationPassA similarly small increment survives a center-held-out challenge
Effect magnitudeInformation onlyDated sources add an internal increment below promotion magnitude
Calendar stabilityFailAdded value is not established across enrollment era
Selective uncertaintyAbstainCoverage is adequate, but uncertainty does not reliably rank error
External PDBP transportFailRecalibration and archive size do not establish transport

What this work contributes

  • A future-only rolling-origin patient representation
  • Separate people for model fitting, calibration, and testing
  • Independent gate verdicts instead of one favorable score
  • Executable stopping decisions under null effects, leakage, and shift

What this work does not claim

  • Individual clinical prognosis
  • Calibrated mechanism probability
  • Treatment selection or treatment benefit
  • Successful external multimodal validation