Most industry benchmarks rely on English-centric data, creating a blind spot for companies deploying AI agents globally. LILT CEO Spence Green warns that while poor translation was once a minor inconvenience, agentic errors in customer-facing roles now carry significant commercial risk. AURORA shifts the focus to native-language performance, utilizing tasks verified by domain experts across sectors like software development, retail, and banking.
In section Releases
LILT Unveils AURORA to Benchmark AI Performance in Non-English Languages
San Francisco-based LILT has launched AURORA, a benchmarking platform designed to test frontier AI models on agentic, multimodal, and socio-cultural tasks beyond English. By moving past traditional translation-heavy metrics, the leaderboard aims to provide enterprises with data on how AI agents function in specific regional and cultural contexts.

The platform evaluates models through four core benchmarks: Multilingual Terminal-bench for localized coding, Multilingual τ³-bench for customer support, Multilingual MultiChallenge for instruction-following, and Multilingual GAIA-v2-LILT for agentic reasoning. Initial analysis from LILT’s PhD-led research team highlights the volatility of model quality, noting that top-performing models shift depending on the language—such as GPT 5.5 excelling in Spanish, Claude Opus 5.5 leading in Japanese, and Muse Spark 1.3 performing best in Serbian.
Comments (0)
No comments yet. Be the first!