Recent advances in Generative Artificial Intelligence (GenAI) and Extended Reality (XR) open new opportunities for immersive learning, but most existing educational XR systems still rely on pre-authored assets and predefined interaction flows, offering poor scalability to heterogeneous and personal study materials. This work addresses that limitation by designing and developing the XR client of an immersive doc-grounded tutoring system, able to turn any static educational document into an interactive, multimodal experience. Developed in Unity for the Meta Quest 3 headset, the system integrates a conversational avatar, a document reader navigable in XR, spatially distributed information panels and a pipeline for the progressive presentation of multimodal artifacts: verbal answers grounded in the document, contextual 2D images and interactive 3D models that can be manipulated hands-free. Every design choice is motivated by the principles of the Cognitive Theory of Multimedia Learning and the Cognitive Load Theory, with the goal of maximizing learning support while containing extraneous cognitive load. Before the user study, the candidate open-weight language models considered as the generative engine were compared through the LLM-as-a-Judge paradigm to select the most reliable one. The system was then validated through a within-subject experimental study with 30 participants, comparing three progressively richer conditions: document only, an embodied conversational tutor, and a spatial tutor with multimodal artifacts. The full spatial condition achieves the best objective learning accuracy (80.0% versus 57.8%), the highest subjective scores for learning and technology acceptance, and the lowest perceived cognitive load, despite its greater functional complexity; usability remains high and comparable across the interactive conditions. The work shows that a multimodal XR tutoring paradigm, grounded in the user's own study materials through real-time AI generation, is technically feasible, pedagogically sound, and perceived as usable and effective by users with varying levels of technological experience.
I recenti progressi nell'Intelligenza Artificiale Generativa (GenAI) e nella realtà estesa (XR) aprono nuove opportunità per l'apprendimento immersivo, ma la maggior parte dei sistemi XR educativi si affida ancora ad asset pre-autorati e a flussi di interazione predefiniti, con scarsa scalabilità rispetto a materiali di studio eterogenei e personali. Questo lavoro affronta tale limite progettando e realizzando il client XR di un sistema di tutoraggio immersivo doc-grounded, capace di trasformare qualsiasi documento educativo statico in un'esperienza interattiva e multimodale. Sviluppato in Unity per il visore Meta Quest 3, il sistema integra un avatar conversazionale, un lettore di documenti navigabile in XR, pannelli informativi distribuiti nello spazio e una pipeline di presentazione progressiva di artefatti multimodali: risposte verbali ancorate al documento, immagini contestuali bidimensionali e modelli tridimensionali interattivi manipolabili a mani libere. Ogni scelta di design è motivata dai principi della Cognitive Theory of Multimedia Learning e della Cognitive Load Theory, con l'obiettivo di massimizzare il supporto all'apprendimento contenendo il carico cognitivo estraneo. Prima della validazione con utenti, i modelli linguistici open-weight candidati come motore generativo sono stati confrontati tramite il paradigma LLM-as-a-Judge per selezionare il più affidabile. Il sistema è stato quindi validato con uno studio sperimentale within-subject su 30 partecipanti, confrontando tre condizioni progressivamente più ricche: solo documento, tutor embodied conversazionale e tutor spaziale con artefatti multimodali. La condizione spaziale completa ottiene la migliore accuratezza di apprendimento oggettivo (80,0% contro 57,8%), i punteggi soggettivi più alti di apprendimento e accettazione tecnologica e il carico cognitivo percepito più basso, nonostante la maggiore complessità funzionale; l'usabilità resta elevata e comparabile tra le condizioni interattive. Il lavoro dimostra che un paradigma di tutoraggio XR multimodale, ancorato ai materiali di studio dell'utente tramite generazione AI in tempo reale, è tecnicamente realizzabile, pedagogicamente fondato e percepito come usabile ed efficace da utenti con diversi livelli di esperienza tecnologica.
Progettazione e Sviluppo di un'Interfaccia XR per il Tutoraggio Educativo Basato su Intelligenza Artificiale Generativa
BORDONI, LEONARDO
2025/2026
Abstract
Recent advances in Generative Artificial Intelligence (GenAI) and Extended Reality (XR) open new opportunities for immersive learning, but most existing educational XR systems still rely on pre-authored assets and predefined interaction flows, offering poor scalability to heterogeneous and personal study materials. This work addresses that limitation by designing and developing the XR client of an immersive doc-grounded tutoring system, able to turn any static educational document into an interactive, multimodal experience. Developed in Unity for the Meta Quest 3 headset, the system integrates a conversational avatar, a document reader navigable in XR, spatially distributed information panels and a pipeline for the progressive presentation of multimodal artifacts: verbal answers grounded in the document, contextual 2D images and interactive 3D models that can be manipulated hands-free. Every design choice is motivated by the principles of the Cognitive Theory of Multimedia Learning and the Cognitive Load Theory, with the goal of maximizing learning support while containing extraneous cognitive load. Before the user study, the candidate open-weight language models considered as the generative engine were compared through the LLM-as-a-Judge paradigm to select the most reliable one. The system was then validated through a within-subject experimental study with 30 participants, comparing three progressively richer conditions: document only, an embodied conversational tutor, and a spatial tutor with multimodal artifacts. The full spatial condition achieves the best objective learning accuracy (80.0% versus 57.8%), the highest subjective scores for learning and technology acceptance, and the lowest perceived cognitive load, despite its greater functional complexity; usability remains high and comparable across the interactive conditions. The work shows that a multimodal XR tutoring paradigm, grounded in the user's own study materials through real-time AI generation, is technically feasible, pedagogically sound, and perceived as usable and effective by users with varying levels of technological experience.| File | Dimensione | Formato | |
|---|---|---|---|
|
Tesi_PDFA.pdf
embargo fino al 11/01/2028
Dimensione
13.1 MB
Formato
Adobe PDF
|
13.1 MB | Adobe PDF |
I documenti in UNITESI sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/20.500.12075/27780