RecruitingRecruiting
Benchmarking Large Language Models Against Tumour Boards for Oncology Treatment Recommendations
NCT07739121 · Assistance Publique - Hôpitaux de Paris
In plain English
Click the button to translate this study into plain language — what it is, who qualifies, and what participation looks like.
Official title
Benchmarking AI for Clinical Oncology decisioNmaking (BEACON): A Prospective, Multicentre, Blinded Evaluation of Frontier Large Language Models Against Multidisciplinary Tumour Board Recommendations in Oncology Treatment Planning
About this study
BEACON is a prospective, multicentre, blinded benchmark using automated, criteria-based scoring. It is built on three design decisions that distinguish it from the existing literature: (i) synthetic, standardised cases remove the record-completeness variability that confounds retrospective comparisons and allow the identical input to be given to every board and every model; (ii) two independent tumour boards per localisation let human-human agreement be measured rather than assumed; and (iii) a guideline matrix, locked a priori, provides an objective anchor applied identically to human and model recommendations.
Reference standard. For each case-domain, a guideline matrix (guideline-recommended / acceptable / unsupported options per case-domain; ESMO, NCCN), locked and time-stamped before data collection, is applied identically to boards and models.
Five decision domains. Every recommendation is decomposed into D1 Intent, D2 Surgery, D3 Radiotherapy, D4 Systemic therapy (class + line), and D5 Work-up \& biomarkers before any comparison.
Eligibility criteria
Inclusion Criteria:
* Synthetic oncology case within one of the five predefined localisations (breast, lung, urological, digestive, gynaecological).
* Complete structured schema: UICC 8th-edition stage, biomarkers, ECOG performance status, comorbidities and a standardised clinical question.
* A clinically answerable treatment-planning question that is mappable to the locked guideline matrix.
Exclusion Criteria:
* Case outside the five predefined localisations.
* Incomplete, internally inconsistent or ambiguous schema.
* Duplicate or near-duplicate of an existing case in the set.
* Question not resolvable by current guidelines.
Study design
Enrollment target: 100 participants
Age groups: adult, older_adult
Timeline
Starts: 2026-05-01
Estimated completion: 2026-10-01
Last updated: 2026-07-31
Interventions
Other: Multidisciplinary tumour boardsOther: Frontier large language models
Primary outcomes
- • Domain-level performance between LLM recommendations and the locked guidelines. (Assessed once at central scoring, after data collection (~October 2026))
Sponsor
Assistance Publique - Hôpitaux de Paris · other
Contacts & investigators
ContactJean-Emmanuel Bibault, MD PhD · contact · jean-emmanuel.bibault@aphp.fr · 01 56 09 34 06
ContactJérôme Lambert, MD PhD · contact · jerome.lambert@u-paris.fr · 0142499742
All locations (1)
Hopital Européen Georges PompidouRecruiting
Paris, France