Study Benchmarks Machine Learning for Antibiotic Stewardship in Pediatric ICUs
A new systematic benchmarking study compares tabular, sequence, and graph-based ML models for predicting antibiotic reduction opportunities in pediatric intensive care units, finding that simpler models often outperform complex ones on calibration.


A systematic benchmarking study published on arXiv compares machine learning architectures for predicting antimicrobial stewardship (AMS) intervention opportunities in pediatric intensive care units (PICUs), with results showing that predictive performance is driven more by target prevalence and dataset characteristics than by model complexity.
The study, led by researchers using the public Paediatric Intensive Care database and a private cohort from the University Children’s Hospital Zurich, Switzerland, addresses a critical gap in clinical AI research. Most prior work on machine learning for antibiotic stewardship has focused on adult populations and static tabular data, leaving pediatric applications underexplored.
Por que importa
Four clinically relevant proxy targets
The research team defined four proxy targets for reducing antibiotic exposure in PICU patients: intravenous-to-oral switching, de-escalation, discontinuation, and short-course therapy. These targets represent real clinical decisions where machine learning could help identify patients ready for reduced antibiotic use.
Under a unified evaluation framework, the researchers compared three model families: traditional tabular models, sequence-based models that process temporal patient data, and graph-based temporal models. The comparison was conducted at multiple temporal resolutions to assess whether finer-grained time windows improve prediction quality.
Contexto
Key findings on model performance
The study found that sequence models improved the precision-recall trade-off over tabular approaches at coarse 24-hour resolution. However, finer temporal modeling provided limited additional benefit, suggesting that daily clinical data may be sufficient for most AMS prediction tasks.
| Key facts | |
|---|---|
| Target clinical decisions | IV-to-oral switch, de-escalation, discontinuation, short-course therapy |
| Model families compared | Tabular, sequence-based, graph-based temporal |
| Data sources | Public Paediatric Intensive Care database + University Children’s Hospital Zurich private cohort |
| Key finding | Simpler tabular models yielded more reliable probability estimates despite lower precision-recall |
A critical finding concerned calibration: while sequence models achieved better precision-recall metrics, they produced less reliable probability estimates. Simpler tabular models, despite lower raw performance, yielded better-calibrated predictions that may be more trustworthy for clinical decision support.
Que sigue
Practical implications for clinical AI development
The results provide practical guidance for building reliable decision support systems for pediatric AMS. The finding that target prevalence and dataset characteristics dominate predictive performance suggests that researchers should prioritize careful target design and data quality over architectural complexity.
The calibration advantage of simpler models is particularly relevant for clinical settings, where poorly calibrated probability estimates could mislead clinicians about the confidence of a model’s recommendation. A model that produces accurate but miscalibrated predictions may lead to inappropriate clinical decisions if practitioners rely on the raw probability scores.
Limitations and next research steps
The study acknowledges several limitations. The proxy targets for antibiotic reduction opportunities may not perfectly capture all clinically meaningful stewardship interventions. The private Zurich cohort introduces potential geographic and practice-pattern biases. Additionally, the study did not assess model performance across different patient subgroups, which is essential for ensuring equitable clinical AI deployment.
Future work should validate these findings across additional pediatric populations and clinical settings, and investigate whether calibration improvements can be achieved without sacrificing precision-recall performance. The research team also plans to explore how these models might integrate into clinical workflows without adding cognitive burden to ICU staff.
Source: arXiv cs.LG – “Benchmarking Machine Learning Architectures for Antimicrobial Stewardship in Pediatric ICUs” (https://arxiv.org/abs/2605.22611)
Source
arXiv cs.LG Publicacion original: 2026-09-16T04:00:00+00:00
Maya Turner
Colaborador editorial.
