The convergence of Extended Reality (XR) and generative Artificial Intelligence is redefining immersive learning, but current XR educational systems often rely on predefined assets and, when present, runtime multimodal generation remains limited to text and static images. This thesis addresses this gap within the framework of ATHENA, a document-grounded XR tutoring system: the user uploads a document, explores it within the virtual environment, and poses vocal questions to an embodied avatar, which responds by drawing solely from the document to mitigate the risk of hallucinations. The central contribution of this work is the design and implementation of a doc-grounded multimodal generation pipeline that, starting from the user's query, coordinately produces a context-grounded response, its relative summary, a contextual image, and an interactive three-dimensional object. The pipeline is implemented as a microservices backend with entirely local inference. It consists of an offline phase, which includes document ingestion, and an online phase (tutoring time), which includes query intent classification, an adaptive retrieval strategy, text generation, meta-prompting for visual and spatial content, text-to-image synthesis, and image-to-3D reconstruction. The choice of the language model to adopt as a backbone was conducted through an LLM-as-a-Judge protocol with component-level benchmarking, prioritizing the trade-off between output accuracy and latency. The second part of the work is dedicated to verifying the educational effectiveness, which was validated through a within-subject study on thirty participants. The study features a nested-condition design (C0 ⊂ C1 ⊂ C2) that isolates the incremental contribution of each support layer. The results highlight how the embodied conversational tutor increases perceived support and system acceptance, but objective learning gain emerges only with the addition of the multimodal and spatial layer. Consequently, the analysis reveals that context-grounded multimodal generation is a promising paradigm for learning in virtual reality.
La convergenza tra Extended Reality (XR) e Intelligenza Artificiale generativa sta ridefinendo l’apprendimento immersivo, ma i sistemi educativi XR attuali dipendono spesso da asset predefiniti e, quando presente, la generazione multimodale a runtime resta circoscritta a testo e immagini statiche. Questa tesi affronta tale lacuna nell’ambito di ATHENA, un sistema di tutoring in XR document-grounded: l’utente carica un documento, lo esplora nell’ambien- te virtuale e rivolge domande vocali ad un avatar incorporato, che risponde attingendo unicamente al documento per contenere il rischio di allucinazioni. Il contributo centrale dell’elaborato è la progettazione e l’implementazione della pipeline di generazione multimodale doc-grounded che, a partire dalla domanda dell’utente, pro- duce in modo coordinato la risposta context-grounded, il relativo riassunto, un’immagine contestuale e un oggetto tridimensionale interattivo. La pipeline è realizzata come backend a microservizi con inferenza interamente locale. Essa si compone di una fase offline, che comprende l’ingestione del documento, e una online (tutoring time), che include la classifica- zione dell’intento della query, una strategia di retrieval adattiva, la generazione testuale, il meta-prompting per i contenuti visivi e spaziali, la sintesi text-to-image e la ricostruzione image-to-3D. La scelta del modello linguistico da adottare come backbone è stata condotta tramite un protocollo LLM-as-a-Judge con benchmark a livello di componente, privilegiando il compromesso tra accuratezza dell’output e latenza. La seconda parte dell’elaborato è dedicata alla verifica dell’efficacia didattica, che è stata validata mediante uno studio within-subject su trenta partecipanti. Lo studio è caratterizzato da un disegno a condizioni annidate (C0 ⊂ C1 ⊂ C2) che isola il contributo incrementale di ciascuno strato di supporto. I risultati evidenziano come il tutor conversazionale incorporato accresca il supporto percepito e l’accettazione del sistema, ma il guadagno di apprendimento oggettivo emerga soltanto con l’aggiunta dello strato multimodale e spaziale. Dunque, l’analisi ha rivelato come la generazione multimodale context-grounded sia un paradigma promettente per l’apprendimento in realtà virtuale.
Intelligenza Artificiale Generativa per l’Apprendimento Immersivo in XR: Sviluppo di un’Architettura Document-Grounded per la Generazione di Contenuti Multimodali e Validazione Sperimentale dell’Efficacia Didattica.
VITALI, JURI
2025/2026
Abstract
The convergence of Extended Reality (XR) and generative Artificial Intelligence is redefining immersive learning, but current XR educational systems often rely on predefined assets and, when present, runtime multimodal generation remains limited to text and static images. This thesis addresses this gap within the framework of ATHENA, a document-grounded XR tutoring system: the user uploads a document, explores it within the virtual environment, and poses vocal questions to an embodied avatar, which responds by drawing solely from the document to mitigate the risk of hallucinations. The central contribution of this work is the design and implementation of a doc-grounded multimodal generation pipeline that, starting from the user's query, coordinately produces a context-grounded response, its relative summary, a contextual image, and an interactive three-dimensional object. The pipeline is implemented as a microservices backend with entirely local inference. It consists of an offline phase, which includes document ingestion, and an online phase (tutoring time), which includes query intent classification, an adaptive retrieval strategy, text generation, meta-prompting for visual and spatial content, text-to-image synthesis, and image-to-3D reconstruction. The choice of the language model to adopt as a backbone was conducted through an LLM-as-a-Judge protocol with component-level benchmarking, prioritizing the trade-off between output accuracy and latency. The second part of the work is dedicated to verifying the educational effectiveness, which was validated through a within-subject study on thirty participants. The study features a nested-condition design (C0 ⊂ C1 ⊂ C2) that isolates the incremental contribution of each support layer. The results highlight how the embodied conversational tutor increases perceived support and system acceptance, but objective learning gain emerges only with the addition of the multimodal and spatial layer. Consequently, the analysis reveals that context-grounded multimodal generation is a promising paradigm for learning in virtual reality.| File | Dimensione | Formato | |
|---|---|---|---|
|
Tesi_Juri_Vitali.pdf
embargo fino al 11/01/2028
Dimensione
4.53 MB
Formato
Adobe PDF
|
4.53 MB | Adobe PDF |
I documenti in UNITESI sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/20.500.12075/28228