Am J Epidemiol. 2026 Aug 25:kwag195. doi: 10.1093/aje/kwag195. Online ahead of print.
ABSTRACT
Short-term recall-derived dietary variables may be relevant to cardiovascular disease, but their classification value after demographic, lifestyle, and clinical markers are observed is unclear in cross-sectional survey data. We analyzed pooled data from 34,388 United States adults in the National Health and Nutrition Examination Survey, 2007-2018, after excluding participants with missing outcome status or current pregnancy. The outcome was a self-reported cardiovascular disease composite defined from coronary heart disease, angina, myocardial infarction, or stroke. We evaluated 3 nested predictor blocks: demographic and lifestyle variables, clinical and laboratory variables, and recall-derived dietary variables. Logistic regression and histogram-based gradient boosting were assessed with 5-fold out-of-fold validation. Adding the clinical block produced the largest discrimination gain (∆ area under the receiver operating characteristic curve 0.031 for logistic regression and 0.032 for gradient boosting; both P <.001). Adding dietary variables after the clinical block produced a 0.002 gain in logistic regression and no material gain in gradient boosting (P = .219). Weighted logistic regression, subgroup, and repeated stochastic imputation sensitivity analyses preserved the same hierarchy. In this setting, recall-derived dietary variables added little classification information once clinical markers were available.
PMID:42639928 | DOI:10.1093/aje/kwag195

