NPJ Cardiovasc Health. 2026 Jul 20. doi: 10.1038/s44325-026-00147-0. Online ahead of print.
ABSTRACT
We aimed to investigate if reinforcement learning (RL) can detect clinical practice variations and if transfer learning (TL) can mitigate resulting distribution shifts across patient sexes and sites. Using data from over 41,000 patients with obstructive coronary artery disease (CAD) in Alberta, Canada from 2009 to 2019, we conducted three experiments: 1) evaluate physician behavior policies across stratified groups using off-policy evaluation (OPE); 2) develop and evaluate RL policies using conservative Q-learning and OPE for each group; and 3) apply TL to improve RL performance on target groups. Female patients had significantly lower expected rewards under their own physician policies compared to males. Applying the physician policies learned from male patients improved rewards for females but remained lower than those for males. RL policies outperformed physician policies. TL with only 10% of target group data effectively mitigated distribution shifts. Residual confounding warrants further research to validate our findings.
PMID:42477048 | DOI:10.1038/s44325-026-00147-0

