This thesis describes the design and development of a software system to support patient screening for the Mayo Clinic GRACE oncology registry, together with the validation of its genetic information extraction pipeline. The work aims to reduce the manual workload of registry coordinators, who currently review medical records, genetic reports, and clinical notes distributed across multiple information systems. The objective is not to replace human evaluation, but to provide a human-in-the-loop tool capable of retrieving, processing, and presenting the evidence relevant to the preliminary assessment of patient eligibility. The developed system is based on a modular and asynchronous pipeline, integrated with a web interface that allows users to launch screening requests and review the generated results. A central component of the GRACE registry concerns genetic risk. For this com- ponent, the system retrieves genetic reports in PDF format, extracts their textual content through direct parsing or Optical Character Recognition (OCR), and uses a language model to transform the extracted text into structured information about relevant genetic mutations. The genetic information extraction pipeline was validated on 101 reports manually an- notated by two researchers with a medical background. Compared with the resulting "gold standard", the system produced an identical set of triplets in 91.1% of the documents. Overall, this work provides a promising software foundation for supporting the review of genetic reports and, more broadly, the patient screening process for oncology registries. The results highlight the potential of combining software pipelines and language models to reduce the manual workload associated with eligibility assessment. At the same time, the sensitive nature of the clinical domain confirms the need to preserve human supervision and to pursue further validation, consolidation, and integration into the operational workflow.
Questa tesi descrive la progettazione e lo sviluppo di un sistema software a supporto dello screening di pazienti per il registro oncologico GRACE della Mayo Clinic, nonché la validazione della relativa pipeline di estrazione delle informazioni genetiche. Il lavoro nasce dall’esigenza di ridurre il carico manuale dei coordinatori del registro, impegnati nell’esame di cartelle cliniche, referti genetici e note cliniche distribuiti tra diversi sistemi informativi. L’obiettivo non è sostituire completamente la valutazione umana, ma fornire uno stru- mento human-in-the-loop capace di recuperare, analizzare e presentare le evidenze rilevanti per la valutazione preliminare dell’eleggibilità dei pazienti. Il sistema sviluppato presenta una pipeline modulare e asincrona, integrata con un’in- terfaccia web per l’avvio dello screening e la consultazione dei risultati. La componente principale del registro GRACE riguarda il rischio genetico, per cui il sistema recupera referti genetici in formato PDF, estrae il testo tramite parsing diretto o Optical Character Recognition (OCR) e utilizza un modello linguistico per produrre informazioni strutturate sulle mutazioni rilevanti. La pipeline di estrazione di informazioni genetiche è stata validata su 101 referti annotati manualmente da due ricercatori con formazione medica. Rispetto al "gold standard" ottenuto, il sistema ha estratto un insieme di triplette identico nel 91,1% dei documenti. Il lavoro costituisce quindi una base applicativa promettente per supportare la revisione dei referti genetici e, più in generale, il processo di screening per registri oncologici. I risultati ottenuti mostrano il potenziale dell’integrazione tra pipeline software e modelli linguistici nel ridurre il carico manuale della valutazione dell’eleggibilità. Allo stesso tempo, la natura sensibile del dominio clinico conferma la necessità di mantenere la supervisione umana e di proseguire con ulteriori fasi di validazione, consolidamento e integrazione nel workflow operativo.
Progettazione e Sviluppo di un Sistema Assistito dall’Intelligenza Artificiale per lo Screening Automatico dell’Eleggibilità dei Pazienti nel Registro Oncologico GRACE
CINGOLANI, FILIPPO
2025/2026
Abstract
This thesis describes the design and development of a software system to support patient screening for the Mayo Clinic GRACE oncology registry, together with the validation of its genetic information extraction pipeline. The work aims to reduce the manual workload of registry coordinators, who currently review medical records, genetic reports, and clinical notes distributed across multiple information systems. The objective is not to replace human evaluation, but to provide a human-in-the-loop tool capable of retrieving, processing, and presenting the evidence relevant to the preliminary assessment of patient eligibility. The developed system is based on a modular and asynchronous pipeline, integrated with a web interface that allows users to launch screening requests and review the generated results. A central component of the GRACE registry concerns genetic risk. For this com- ponent, the system retrieves genetic reports in PDF format, extracts their textual content through direct parsing or Optical Character Recognition (OCR), and uses a language model to transform the extracted text into structured information about relevant genetic mutations. The genetic information extraction pipeline was validated on 101 reports manually an- notated by two researchers with a medical background. Compared with the resulting "gold standard", the system produced an identical set of triplets in 91.1% of the documents. Overall, this work provides a promising software foundation for supporting the review of genetic reports and, more broadly, the patient screening process for oncology registries. The results highlight the potential of combining software pipelines and language models to reduce the manual workload associated with eligibility assessment. At the same time, the sensitive nature of the clinical domain confirms the need to preserve human supervision and to pursue further validation, consolidation, and integration into the operational workflow.| File | Dimensione | Formato | |
|---|---|---|---|
|
Tesi_Cingolani_AM_rev2-a.pdf
embargo fino al 11/01/2028
Dimensione
1.75 MB
Formato
Adobe PDF
|
1.75 MB | Adobe PDF |
I documenti in UNITESI sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/20.500.12075/27783