A status card for VitalMatch's severe-PGD@72h classifier: how it compares to published donor-lung scores on a held-out test set, what it can and can't yet claim, and the caveats that keep the numbers honest. Research use only — not a validated clinical tool.
Source: internal held-out evaluation (seed-grouped, donor-level split) on the STAR bilateral-lung cohort. Published scores adapted to available STAR fields.
0.68
The clinical model's robust discrimination is an AUROC of 0.68 (cross-seed median) for severe PGD@72h — modest, but it beats the published donor-lung scores, which land between chance and ~0.64 on this modern cohort.
01 · The test set
Small, honest, donor-level split
251
Held-out recipients (donor-grouped)
22
PGD events — 8.8% prevalence
±0.1
approx AUROC uncertainty at this event count
With only 22 events, every AUROC here carries a wide confidence interval — differences of a few hundredths are not decisive. This card reports cross-seed medians, not a single lucky split.
02 · Benchmark
Beating published scores that are weak on modern data
AUROC for severe PGD@72h: published donor-lung scores vs the VitalMatch clinical model, same held-out cohort. Published scores — built on older eras and adapted to STAR fields — sit close to chance.
AUROC (0.50 = chance)
Bars proportional from 0; values labelled. Published scores shown at the upper end of their adaptation range.
03 · What we don't claim
The honest caveats
A random forest reaches AUROC 0.92 cross-seed on the same data — but with a large train–test gap on 251 cases, that number is almost certainly optimistic and is not yet validated; we lead with the linear model's 0.68. Imaging adds no incremental value yet — in late-fusion tests the CT branch is effectively ignored (the clinical panel carries the signal), so the current headline is clinical-only. And a strong AUROC is not calibration: decision-curve and calibration analysis, and external validation, are the next milestones before any clinical claim.