A Safety-Supervised Dual-Agent Reinforcement Learning Framework for Pediatric Automated Insulin Delivery
Lops, G.; De Cicco, L.; Mascolo, S.
Abstract
This study presents a safety-supervised dual-agent reinforcement-learning
framework for pediatric automated insulin delivery. The architecture combines
a primary Maskable Proximal Policy Optimization controller with predictive
correction, learned insulin-command modulation, and deterministic runtime
constraints. The framework was evaluated in silico using ten virtual
pediatric subjects, five independent training seeds, and 200 matched
stochastic 24-h scenarios per subject and seed.
Relative to an information-matched single-agent Maskable PPO baseline, the
dual-agent framework increased time in range (TIR) in all ten subjects,
raising the cohort mean from 42.86% to 69.96% (paired difference:
27.10 percentage points; 95% CI: [20.80,33.40]; p<0.001).
Cohort-average time below range decreased from 25.56% to 11.99%,
while time above range decreased from 31.58% to 18.05%. The Blood
Glucose Risk Index decreased from 22.31 to 8.82, corresponding to a
$60.5\%$ relative reduction (p=0.007), and inter-subject TIR variability
decreased from 33.90% to 13.11. Component-wise ablation showed
complementary contributions from learned modulation and deterministic
runtime constraints, while subject-specific analyses demonstrated adaptive
responses across heterogeneous glycemic profiles.
Overall, the results provide simulation-based evidence that hierarchical
integration of learned and deterministic supervision can improve glycemic
regulation and consistency across the evaluated pediatric profiles and
stochastic meal scenarios. The proposed framework offers a modular basis for further
investigation of safety-aware reinforcement learning in automated insulin
delivery.