Paper 1 | An evidence-quality spine
Validity Gates Expose Limits of Longitudinal Multimodal Parkinson Disease Prediction
A model does not earn credibility from one favorable score. We ask seven separate questions: Was the future hidden? Were people and sites separated? Is the gain large enough to matter? Does it survive calendar change? Can uncertainty identify safer forecasts? Does the result travel to another cohort?
The result at a glance
Seven questions produce seven verdicts, not one headline score
Each gate protects a different claim. Passing chronology cannot rescue failed transport, and correct interval coverage cannot substitute for useful error ranking.
The scientific output is the decision boundary. Dated multimodal history contains a reproducible internal increment, but the effect is clinically tiny, temporally fragile, and not externally transportable in the available export. The protocol therefore stops the claim before deployment language begins.
Study design
Every forecast begins at a dated cutoff
Imagine pausing one patient's record on a particular day. Only measurements already available at that moment may describe the state. Their age and absence remain visible; the future motor examination, 9 to 15 months later, stays hidden until evaluation.
Why ordinary validation is insufficient
A repeated clinical history is not a flat table
The same person appears at multiple visits. A scan may be months older than the examination it accompanies. Missingness may encode a study protocol rather than biology. Medication state changes motor scores, and a new enrollment era can change both who is measured and how sources are collected.
These are different scientific risks. The benchmark therefore assigns each one a distinct population, uncertainty interval, and promotion rule.
| Analysis | Scale | Question it answers |
|---|---|---|
| Rolling-origin benchmark | 1,058 people; 5,404 cutoffs | Do dated sources add internal information when every person stays in one fold? |
| Strict-calendar challenge | 442 later-era participants | Does the frozen relation survive a later enrollment and observation regime? |
| Site-disjoint challenge | 951 people across 52 sites | Is the increment dependent on acquisition center? |
| Gate stress test | 8,000 controlled experiments | Does the protocol suppress false promotion and detect leakage or shift? |
| PDBP clinical transport | 77 calibration; 156 untouched test participants | Does a frozen clinical model travel to an independent cohort? |
Internal result
The signal is measurable; the improvement is not clinically persuasive
A statistically stable improvement can still be practically negligible. The 0.060 point MAE reduction does not establish individual prognosis, patient-management value, or clinical-trial enrichment.
Uncertainty and stress testing
An interval may cover well and still fail to identify safer predictions
Coverage asks whether the future falls inside an interval. Selective prediction asks a harder question: when the model says it is more certain, is it actually less wrong?
External transport
A large archive is not the same as a compatible cohort
PDBP contains valuable Parkinson disease data, but a transport claim requires the same diagnosis, modalities, timing, endpoint, and eligible participants to intersect. File presence alone is not validation.
Claim contract
What the evidence permits us to say
| Gate | Verdict | Permitted statement |
|---|---|---|
| Future-only chronology | Pass | Every predictor precedes its endpoint |
| Participant separation | Pass | The internal signal is not produced by repeated-person leakage |
| Site separation | Pass | A similarly small increment survives a center-held-out challenge |
| Effect magnitude | Information only | Dated sources add an internal increment below promotion magnitude |
| Calendar stability | Fail | Added value is not established across enrollment era |
| Selective uncertainty | Abstain | Coverage is adequate, but uncertainty does not reliably rank error |
| External PDBP transport | Fail | Recalibration and archive size do not establish transport |
What this work contributes
- A future-only rolling-origin patient representation
- Separate people for model fitting, calibration, and testing
- Independent gate verdicts instead of one favorable score
- Executable stopping decisions under null effects, leakage, and shift
What this work does not claim
- Individual clinical prognosis
- Calibrated mechanism probability
- Treatment selection or treatment benefit
- Successful external multimodal validation