Background
Open-air heritage sites delivering live guided tours and craft demonstrations face persistent multilingual accessibility barriers. At Āraiši Ezerpils Archaeological Park, expert guides and craft demonstrators have deep practical knowledge of 9th-10th century Latgalian construction, textile, metalworking, and daily-life practices, but linguistic capability is largely Latvian with some English or German, creating gaps for international visitors and excluding access for visitors speaking other European languages. Conventional screen-based translation applications proved unsuitable for the open-air, demonstration-heavy context: visitors focused on phone screens reading translated text miss the gesture, tool manipulation, and material handling that carry essential information beyond pure linguistic content. The VAARHeT sub-project Pilot 3, funded through the EU Horizon Europe VOXReality Open Call cascade mechanism (Grant agreement 101070521), proposed addressing this through an AR translation agent combining VOXReality Automatic Speech Recognition and Neural Machine Translation, with translated text delivered via mobile phone subtitle overlay or peripheral-vision AR wearable display, allowing visitors to keep visual attention on the live demonstration.
Technical architecture
XR Ireland, Nuwa's sister brand, developed an Android application running on Samsung Galaxy Note10+ 5G smartphones, optionally paired by Bluetooth to ActiveLook micro-OLED AR wearable glasses for hands-free peripheral-vision text display. Source and target languages were configured through dropdown selectors offering German, English, and Latvian, covering primary visitor demographics at Latvian archaeological sites serving Central European tourism markets. The architecture distributed work across edge and cloud:
- VOXReality Automatic Speech Recognition ran on-device, transcribing tour guide speech in real time
- VOXReality Neural Machine Translation ran on NVIDIA A100 GPU cloud infrastructure, processing source-language text through sequence-to-sequence models trained on parallel corpora including general domain text and cultural heritage terminology, returning translated text to the mobile application
- The mobile UI rendered camera passthrough with subtitle text overlaid at the bottom of screen in the established video-subtitle convention; AR glasses rendered translated text in upper peripheral vision optimised for the micro-OLED display
A push-to-talk button initiated fixed-duration recording segments transmitted for translation as discrete chunks rather than continuous streaming, introducing an interaction pattern requiring visitor timing discipline. End-to-end latency from push-to-talk activation through ASR, cloud transmission, NMT processing, return transmission, and text rendering achieved median 2076ms (mean 2256ms, 95th percentile 3203ms), the highest latency of the three VAARHeT pilots but still meeting the project sub-2500ms KPI in 90% of cases, with 91.9% participant rating as acceptable or very acceptable speed.
Validation methodology
Validation engaged 37 participants at Āraiši Ezerpils across 14-16 July 2025, with inclusion criteria matching other VAARHeT pilots plus particular emphasis on bilingual or multilingual capability enabling informed assessment of translation quality rather than purely interface usability. Language pair testing was expanded opportunistically beyond the primary German-English specification:
- German to English: 19 tests
- English to Latvian: 16 tests
- German to Latvian: 3 tests
- Latvian to English: 1 test
- Latvian to German: 1 test
Cordula Hansen of Technical Art Services designed the test scenarios across wearable donning and comfort, mobile application startup and Bluetooth pairing, language selection configuration, and translated text reading comprehension on both displays. Technical failure definition for this pilot extended beyond software malfunction to include ergonomic limitations: inability to wear AR glasses due to prescription eyewear conflicts, text illegibility from display brightness or resolution constraints, and physical discomfort preventing sustained wearable use, recognising that hardware usability proved as critical as software functionality for deployment viability. Post-test surveys assessed first impressions, Net Promoter Score, hardware comfort, translation accuracy perceptions, text legibility on both display modalities, response speed, and technostress symptoms including eye strain.
Hardware usability outcomes
Mobile and AR glasses produced dramatically different usability profiles. Mobile-side text reading achieved 94.6% completion without help versus 83.8% for AR glasses, with 5.4% technical failure from illegibility on the wearable. Wearable donning achieved 89.2% without help, but mobile application startup and Bluetooth pairing dropped to 64.9% completion without help, with 32.4% needing tester assistance to complete the multi-step configuration. Source language selection reached 75.7% without help, target 89.2%, suggesting asymmetric difficulty where the initial language selector interaction proved more challenging than the subsequent repetition of the same UI pattern.
The ActiveLook glasses showed only 27% strong agreement on comfort, 16.2% strong disagreement, and combined 29.7% negative perception. Text legibility on the wearable was 40.5% positive against 45.9% negative, while mobile text legibility was 89.1% positive against 2.7% negative. Free-text feedback identified four hardware constraints: glasses weight on the nose bridge during extended wear, incompatibility with prescription eyewear, micro-OLED resolution and outdoor brightness inadequate for legibility, and focal distance mismatch between real-world demonstration observation and near-eye display reading creating eye strain through continuous accommodation adjustment. Multiple participants explicitly preferred the mobile display despite initially expecting AR glasses would prove superior, with several defaulting to the phone screen even when glasses were operational. Verbatim feedback such as "glasses are awkward, but the application is good" and "good idea but not working as expected" captured the consistent reception pattern.
Translation quality variance across language pairs
Translation quality varied dramatically by language pair, in direct proportion to parallel-corpus availability for each language within the VOXReality NMT component. German to English, a high-resource pair with decades of machine translation research investment, was rated reliable and comprehensible by participants, with occasional non-idiomatic phrasing but semantically correct output. English to Latvian and other Latvian-involving pairs degraded severely, with participants describing output as "very poor" or "comical", including non-existent word inventions combining morphemes incorrectly, repetitive phrase generation, grammatical violations of Latvian linguistic rules, and semantic failures where the target language conveyed meaning unrelated or opposite to source speech. Structured assessment showed only 8.1% strong agreement that translations felt accurate and reliable, 40.5% agreement, 18.9% neutral, 16.2% disagreement, and 16.2% strong disagreement.
The Latvian quality problem was particularly significant given the museum's location and primary domestic visitor base. Latvian to English or Latvian to German translation would enable international visitors to attend regular scheduled Latvian-language tours rather than requiring special English-language guide booking, yet output quality prevented practical deployment despite this representing the highest-value use case for Āraiši Ezerpils operationally. The variance demonstrates a critical European AI development priority: commercial NMT platforms optimised for high-resource languages prove insufficient for heritage sectors serving linguistically diverse populations, requiring dedicated investment in parallel corpus development, domain-specific terminology training, and continuous quality improvement that general commercial translation services do not prioritise for specialised cultural heritage vocabulary.
Net Promoter Score and strategic recommendations
Net Promoter Score was -14 (12 promoters, 8 passives, 17 detractors out of 37), the only VAARHeT pilot to receive a net negative recommendation likelihood and positioning the translation agent in "needs significant improvement" territory. The detractor population of 46% was concentrated among participants experiencing Latvian language pair translations and among users attempting sustained AR glasses use. Participants testing German-English on mobile display gave markedly more positive feedback, suggesting that successful language pair combined with usable display hardware could achieve acceptable satisfaction comparable to the VR Site Augmentation pilot. Appropriateness to museum context reached 64.8% positive (29.7% strong agreement, 35.1% agreement), substantially below VR Site Augmentation's 97.4% but above Welcome Avatar's 55.3%.
The validation generated three strategic recommendations. First, AR wearable glasses are not currently recommended for cultural heritage translation: commercially available European hardware proves inadequate for sustained text reading across comfort, legibility, and acceptance dimensions, with mobile phone displays providing superior experience without additional procurement burden. Second, minority language translation quality must reach acceptable accuracy before any Latvian heritage site deployment, requiring additional training investment, domain-specific corpus development, and heritage terminology validation. Third, commercial differentiation is questionable given established alternatives from Microsoft, Google, and Meta wearables addressing similar use cases with broader language support and mature quality assurance. For Culturama Platform development, this points toward integration partnerships with established translation API providers rather than proprietary translation development, with Nuwa concentrating on heritage-specific capabilities including curator-validated archaeological terminology, multilingual content management workflows, and integration with European heritage aggregators like Europeana. As with Pilot 1's Welcome Avatar, the honest finding informed strategy as much as a positive result would: validating where Nuwa should not invest is equally valuable to confirming where it should.
