Safety-aware AI for NSCLC trial pre-screening: a comparative proof-of-concept study of rule-based, single-agent, and multi-agent approaches.
Safety Agent gating reduced false inclusion in non-small cell lung cancer trial pre-screening.
Safety-aware AI for NSCLC trial pre-screening: a comparative proof-of-concept study of rule-based, single-agent, and multi-agent approaches.
Non-small cell lung cancer (NSCLC) trials often involve complex biomarker-driven and line-specific eligibility criteria, making pre-screening labor-intensive and error-prone.
To compare safety-aware AI configurations for NSCLC trial pre-screening, with explicit focus on reducing false inclusion rates and supporting uncertainty-aware abstention under evidence limitations.
We performed a retrospective proof-of-concept evaluation using interventional NSCLC trial eligibility text and structured synthetic patient summaries covering variations in stage, biomarker status, prior therapy, ECOG performance status, and exclusion-relevant factors.
Safety Agent gating reduced incorrect patient inclusion to 1.16% for non-small cell lung cancer trials.
The single-agent LLM (GPT-4o-mini) achieved the highest observed accuracy (0.74, 95% CI 0.66-0.81), with no observed false inclusion among 86 gold-standard ineligible pairs (95% CI 0.0-4.3%) and an uncertain rate of 7.5%.
In the TCGA-derived stress test with missing biomarker and ECOG data, all evaluated systems showed zero observed false inclusion.
In this proof-of-concept evaluation under controlled synthetic-input conditions, conservative safety gating substantially reduced false inclusion in rule-based pre-screening, and all LLM-based configurations achieved zero or near-zero false inclusion on the structured synthetic cohort.
These findings support explicit false inclusion monitoring, uncertainty-aware abstention, and multi-dimensional evaluation as useful design principles for AI-assisted oncology trial pre-screening.