Background
Āraiši Ezerpils Archaeological Park in Latvia, home to Europe's only reconstructed 9th-10th century fortified lake settlement, faced operational pressures common to open-air heritage institutions: limited permanent staff bottlenecked at reception during summer peak, seasonal demand concentrations leaving shoulder periods with minimal interpretive support, multilingual accessibility gaps for international visitors, and tight budgets ruling out digital kiosks or expanded staffing. The VAARHeT sub-project, funded through the EU Horizon Europe VOXReality Open Call cascade mechanism (Grant agreement 101070521), proposed addressing these challenges through a voice-activated mobile AR welcome avatar deployable on visitors' own smartphones without additional infrastructure investment.
Technical architecture
XR Ireland, Nuwa's sister brand, led technical development integrating three VOXReality AI components within a Unity-based Android application running on Samsung Galaxy Note10+ 5G devices. ARCore plane detection anchored a 3D avatar on the museum floor at a visitor-selected position. A push-to-talk button captured voice input, with the avatar replying via text panels rather than spoken audio. The three AI components ran across a hybrid on-device and cloud architecture:
- VOXReality Automatic Speech Recognition ran on-device, transcribing speech to text in real time
- VOXReality Intent Classifier ran in cloud, matching natural language queries against eight authorised intents covering facility locations, ticket pricing, event schedules, attraction listings, directions, opening hours, safety restrictions, and general context queries
- VOXReality Dialogue System generated responses through Retrieval Augmented Generation against a curator-validated knowledge base of museum documentation, FAQs, and historical context supplied by Āraiši Ezerpils staff
All cloud inference ran on NVIDIA A10G GPU infrastructure in EU-compliant German cloud facilities. Text-to-speech audio output was cut from scope for budget and timeline reasons, with visitors receiving text responses from a visual 3D avatar, a combination that proved a significant source of disappointment in testing.
Validation methodology
Usability testing ran 14-16 July 2025 at the park with 39 recruited participants aged 25-75. The cohort largely matched the museum's primary visitor demographic: predominantly female (59%), aged 30-50 with professional occupations, Latvian native speakers with good English comprehension, daily mobile users, with around half having prior digital wearable or VR exposure and one-third with AR application experience. Over 75% had used conversational chatbots, more than 50% reporting daily chatbot interaction. Cordula Hansen of Technical Art Services designed the validation methodology following standard mixed-methods user experience research practice, with ethical approval obtained through Maynooth University. Test sessions used moderated in-person protocols capturing task completion (successful, with help, abandoned, technical failure), Net Promoter Score, and Likert-scale assessments of avatar appropriateness, voice interaction naturalness, information relevance, and accuracy perception.
Task completion outcomes
Application launch achieved 94.9% completion (5.1% needed help). Language selection reached 100%. AR avatar placement proved the most consistent friction point at 74.4% without help, with 25.6% requiring tester guidance to complete the ARCore floor-scan interaction, suggesting spatial interaction patterns were less intuitive than conventional mobile UI elements. Information retrieval varied substantially by query type: ticket prices 89.7%, attraction listings 84.6%, bathroom locations 79.5%, directions 76.9%, and event schedules just 64.1%, with time-sensitive accuracy proving more demanding than static facility details. Technical failures (speech recognition errors, intent classification ambiguity, RAG hallucination, or network timeouts) affected 2.6 to 10.3% of attempts across query types. Median end-to-end response latency was 1766ms with 95th percentile at 1898ms, comfortably under the project's 2500ms KPI; 92.1% of participants rated response speed as acceptable or very acceptable, exceeding the 85% KPI threshold.
User experience and avatar reception
Polarised reception revealed tension between modality appreciation and content concerns. 78.9% agreed the avatar was a positive addition to the museum experience, but only 55.3% considered it appropriate to the museum context, with 31.6% disagreeing on cultural fit. The contemporary avatar styling (modern clothing, blue hair) drew consistent criticism: visitors expected historically accurate 9th-10th century Latgalian costume aligning with the museum's interpretive mission. The text-only response modality drew particular disappointment, with many participants expecting spoken audio from a visual 3D representation. Latvian language absence was a major access issue, excluding the domestic visitor population in favour of English-speaking international tourists.
Information accuracy was the critical failure mode. Only 47.3% agreed responses were accurate, with roughly one in four responses estimated to contain factual errors or hallucinations. The Nielsen severity framework rated hallucinated responses as a severity-4 usability catastrophe requiring mandatory resolution before acceptable deployment. Voice interaction naturalness reached 63.2% agreement; efficiency 73.7%; relevance 60.5%. The Net Promoter Score landed at 16 (16 promoters, 12 passives, 10 detractors), positioning the avatar in "needs improvement" territory rather than the "great" or "excellent" range typical of successful museum technology deployments.
Critical lessons for heritage AI
The validation generated four lessons directly informing Nuwa's Culturama Platform roadmap. First, theoretical content delivery in cultural heritage benefits from conventional digital modalities (FAQs, simple chatbots, digital signage) rather than immersive XR, with the avatar's spatial AR representation creating cognitive load disproportionate to the information-retrieval task. Second, heritage institutions require substantially higher AI accuracy thresholds than commercial chatbots, with curator validation, guardrails preventing speculative responses, explicit uncertainty acknowledgement, and source attribution proving mandatory rather than optional. Third, minority and regional language support is essential for European heritage applications, with English-only deployment contradicting the cultural mission to serve regional populations as primary constituencies. Fourth, voice interaction provides genuine convenience value, but speech recognition failures, intent classification errors, or factual response inaccuracies erode trust more severely than equivalent failures in typed interfaces, where users attribute errors to their own input rather than system capability.
The team concluded that a simple chatbot without AR representation would likely serve the routine information delivery use case more efficiently, with XR investment better directed at experiential applications such as virtual archaeological site exploration, temporal reconstruction, and interactive demonstration of historical construction and craft techniques, where immersive technology provides unique value impossible to replicate through conventional media. This finding aligns with parallel validation in Nuwa's XRisis humanitarian training project, where AI dialogue avatars also delivered limited incremental value for theoretical knowledge transfer compared with substantial gains for situated soft-skills practice.
