Background
Open-air archaeological museums face fundamental interpretation challenges that conventional guided tours and static exhibits struggle to address effectively. At Āraiši Ezerpils Archaeological Park, visitors encounter faithful reconstructions of 9th-10th century Latgalian lake settlement dwellings built using authentic archaeological evidence, yet the architectural complexity of Iron Age building techniques (foundation pile systems anchoring buildings into lakebeds, log wall construction without metal fasteners, complex roof framing) is largely invisible to visitors viewing completed reconstructions from exterior vantage points. Live tour guides explain effectively but scale poorly under staffing limitations, linguistic barriers, and Northern European winter conditions that prevent comfortable outdoor tours despite institutional interest in year-round engagement. The VAARHeT sub-project Pilot 2, funded through the EU Horizon Europe VOXReality Open Call cascade mechanism (Grant agreement 101070521), proposed addressing these challenges through voice-activated VR site augmentation, deployed on Meta Quest 3 standalone headsets in climate-controlled indoor exhibition space, enabling visitors to explore a 3D reconstruction of the Chieftain's house at 1:1 scale through natural speech commands.
VR experience design
XR Ireland, Nuwa's sister brand, built the experience in Unity targeting the Meta Quest 3 platform. The 3D reconstruction modelled six discrete interactive building components, each validated by Āraiši Ezerpils heritage experts for archaeological accuracy: floor foundation and pile system, log wall assemblies with notching and joining details, door and window aperture framing, interior wooden bench installations, clay oven and chimney integration, and roof framing with weatherproofing. Visitors triggered educational content by speaking naturally ("show me how the floor was built", "explode the walls", "zoom in on the oven"), with VOXReality components handling the speech pipeline:
- VOXReality Automatic Speech Recognition ran on-device within the Quest 3 Android-based OS
- VOXReality Intent Classifier ran in cloud on NVIDIA A100 GPU infrastructure, mapping natural language queries against six authorised intents (floor, walls, doors, benches, oven, roof)
- Pre-scripted educational content was triggered by recognised intents, deliberately chosen over AI-generated narrative to prevent hallucination or speculative interpretation that could undermine educational integrity, with precise synchronisation between voice events, camera animations, 3D model transformations, and explanatory text
Four of the six content categories (walls, benches, oven, roof) had visual text labels prompting visitors about available queries. Two categories (floor, doors) remained intentionally unlabelled to assess content discoverability in pure voice-only interaction. Median end-to-end latency was 1960ms from question completion to triggered animation, with 89.7% of participants rating response speed as acceptable or very acceptable.
Validation results
Validation testing ran 14-16 July 2025 with 38 participants (one of the 39-person cohort opted out due to prior motion sensitivity concerns). Headset donning achieved 86.8% completion without help, and 89.5% successfully triggered at least one animation event through voice interaction. Individual content category completion revealed the discoverability problem clearly: visually-labelled content scored 81-87% without assistance (walls 84.2%, benches 81.6%, oven 86.8%, roof 86.8%), while the unlabelled categories dropped to 47.4% for floor and 50% for doors, with 21-26% of participants abandoning attempts before discovering the content. Approximately 25% of sessions experienced unintentional content triggering from background conversation or thinking-aloud speech, an inherent constraint of continuous listening voice interfaces that cannot reliably distinguish intentional commands from casual remarks. Zero participants reported cybersickness severe enough to terminate sessions, validating the conservative scene design with stationary viewer position, smooth camera transitions, and static environmental elements.
Educational value and user experience
VR Site Augmentation produced the strongest validation outcomes of the three VAARHeT pilots. Net Promoter Score reached 61 (25 promoters, 13 passives, 1 detractor), positioning the experience in the "great" category with a 45-point differential over the Welcome Avatar pilot. 66.7% of participants strongly agreed the VR experience was a positive museum addition, with 97.4% positive or neutral. Appropriateness to museum context achieved 97.4% positive with only a single strong dissenter. Information accuracy perception reached 89.7% positive agreement (46.2% strong, 43.6% agree), validating the pre-scripted content approach versus the RAG-generated responses that downgraded Welcome Avatar's accuracy ratings.
Participant feedback consistently emphasised content quality, immersive presence sensation ("feels like I'm there"), and calm atmospheric presentation aligned with contemplative learning rather than voice mechanics as the primary value drivers. Critical feedback focused on interaction modality limitations rather than content quality: restricted access from a limited intent set, uncertainty about system listening state ("I wasn't sure if there was anyone there to hear me"), language limitation ("it would be better in my native language"), and occasional wrong content triggering. Notably absent from criticism were concerns about educational accuracy, visual quality, or cultural appropriateness that had dominated Welcome Avatar feedback.
Behavioural insights and voice-only friction
Direct observation captured interaction patterns that structured surveys would not. Approximately 40% of participants attempted Latvian queries despite English-only deployment clearly explained during the pre-test briefing, reinforcing minority language support as a critical success factor rather than optional enhancement. Many visitors asked questions beyond the six facilitated intents, attempting open-ended architectural queries or cultural practice explorations that the intent classifier could not map, revealing an expectation that voice interaction implies comprehensive conversational capability when the implementation supported only narrow pre-defined categories. Visitors frequently looked around the environment passively without speaking, indicating that voice interaction introduces an unfamiliar engagement paradigm requiring explicit interface prompting to transform passive viewers into active questioners. Multiple participants attempted hand gestures for interaction despite VR controllers remaining inactive and extensive pre-test briefing emphasising voice-only modality, suggesting voice-only interaction is sufficiently novel that users default to familiar manipulation patterns until repeated experience establishes new mental models. Social self-consciousness about speaking aloud in a semi-public testing environment ("Am I too loud?") affected several participants.
Strategic implications
The validation generated three design principles informing Nuwa's Culturama Platform roadmap. First, pure voice-only interaction creates discoverability problems where visitors cannot determine what content exists without extensive experimentation that many users abandon. Hybrid voice-visual interfaces combining visual menu representations or content availability indicators with voice activation mechanics preserve discoverability while leveraging voice convenience for hands-free selection. The 30-40 percentage point gap between labelled and unlabelled content completion quantifies this trade-off precisely. Second, accidental triggering from conversational speech in continuous listening mode warrants push-to-talk activation despite the minor friction it adds, particularly in indoor spaces where multiple visitors might use the application simultaneously. Third, content quality and archaeological accuracy substantially exceeded interaction sophistication in importance for visitor satisfaction, validating development investment priorities concentrating on curator-validated knowledge, expert-reviewed 3D reconstructions, and authentic historical representation rather than pursuing cutting-edge AI generation that could compromise factual correctness.
Of the three VAARHeT pilots, this is the one Āraiši Ezerpils chose to install permanently. The museum's plan to maintain the VR experience at the park for long-term visitor feedback collection demonstrates the strongest path from research validation to operational integration of any VAARHeT pilot. Technology Readiness Level 7 was achieved through operational environment validation with authentic visitor populations, positioning the application for continued commercial development. For Culturama Platform, the lesson is clear: heritage XR investment is best directed at experiential applications where spatial presence and immersive content delivery provide unique capabilities impossible to replicate through conventional media, with voice interaction positioned as accessibility enhancement rather than core differentiator.
