A safety-supervised dual-agent reinforcement learning framework for pediatric automated insulin delivery
Lops, G.
Abstract
This study presents a safety-supervised dual-agent reinforcement-learning framework for pediatric automated insulin delivery. The architecture combines a primary Maskable Proximal Policy Optimization controller with predictive correction, learned insulin-command modulation, and deterministic runtime constraints. The framework was evaluated in silico using ten virtual pediatric subjects, five independent training seeds, and 200 matched stochastic 24-h scenarios per subject and seed. Relative to an information-matched single-agent Maskable PPO baseline, the dual-agent framework increased time in range (TIR) in all ten subjects, raising the cohort mean from 42.86% to 69.96% (paired difference: 27.10 percentage points; 95% CI: [20.80,33.40]; p < .001). Cohort-average time below range decreased from 25.56% to 11.99%, while time above range decreased from 31.58% to 18.05%. The Blood Glucose Risk Index decreased from 22.31 to 8.82, corresponding to a 60.5% relative reduction (p=.007), and inter-subject TIR variability decreased from 33.90% to 13.11%. Component-wise ablation showed complementary contributions from learned modulation and deterministic runtime constraints, while subject-specific analyses demonstrated adaptive responses across heterogeneous glycemic profiles. Overall, the results provide simulation-based evidence that hierarchical integration of learned and deterministic supervision can improve glycemic regulation and consistency across the evaluated pediatric profiles and stochastic meal scenarios. The proposed framework offers a modular basis for further investigation of safety-aware reinforcement learning in automated insulin delivery.