Age Ageing. 2026 Sep 4;55(9):afag262. doi: 10.1093/ageing/afag262.
ABSTRACT
BACKGROUND: The frailty index (FI) assumes that deficits are interchangeable; the type of individual items matters less than the total count. Machine learning is increasingly used to create shorter FIs, often motivated by perceived barriers to collecting many items in clinical settings despite availability of electronic health records. We explored whether FI item selection matters when predicting mortality in adults living with cardiovascular disease (CVD).
METHODS: We analysed 3669 adults living with CVD from the National Health and Nutrition Examination Survey (1999-2018) with a 46-item FI. Cox and random survival forest models predicted all-cause and CVD-specific mortality using individual items or composite FI scores. Items were ranked by permutation importance. At each item count k (1-46), we compared the top k items as individual predictors, the top k items combined into a single FI and 200 FIs each built from randomly selected k items.
RESULTS: Models using 46 individual items outperformed composite FI approaches for both outcomes (C-index: 0.71 [0.69-0.74] vs 0.64-0.67). Importance-ranked FI performance peaked at ~10 items then declined. Randomly built FIs improved steadily, converging with the importance-ranked composites by ~35 items (all-cause) and 20 items (CVD). Analyses using imputation produced consistent results.
CONCLUSIONS: When enough items are included, which specific items make up the FI has negligible influence on prediction. Combining items into a single score discards some prognostic information, but this trade-off maintains the generalisability that makes the FI practical. Short importance-ranked FIs benefit from avoiding dilution rather than capturing an optimal item set, and any such selection is outcome-specific and sample-dependent.
PMID:42696642 | DOI:10.1093/ageing/afag262

