European Multilingual Evaluation
A rigorous benchmark evaluating LLM quality beyond English, across all 24 official EU languages in professional and institutional contexts.
2,800+
Planned tasks
24
EU languages
4
Task genres
2026
Release
Leaderboard coming 2026
First results will be published alongside the initial release. Contact the research team to participate in the pilot evaluation.
Languages covered
All 24 official EU languages · native-authored tasks
Related research
Published work this benchmark builds on — and the gap it addresses
Singh et al. (Cohere For AI), 2024 · arXiv:2412.03304
Shows that translated benchmarks carry Western-centric cultural bias and that model rankings shift across languages — motivating native-authored, not translated, EU tasks.
Bandarkar et al. (Meta AI), 2023 · arXiv:2308.16884
Documents substantial comprehension gaps between high- and lower-resource languages, including several official EU languages.
openGPT-X, 2024
Evaluates models across 21 European languages — but with machine-translated versions of English benchmarks (ARC, MMLU, GSM8K), the gap KAROKAN-LANG's native-language tasks address.
Ahuja et al., 2023 · arXiv:2303.12528
16 datasets across 70 languages: generative model performance drops sharply outside high-resource languages.
Nielsen et al. (grew out of ScandEval, arXiv:2304.00906)
Continuously updated NLU/NLG evaluation for European languages — complementary coverage focused on shorter-form tasks rather than professional workflows.
About
KAROKAN-LANG systematically evaluates large language models across all 24 official EU languages using tasks sourced from real professional and administrative settings. It targets the performance gaps hidden by English-only benchmarks and reveals how model quality degrades — or holds — as context shifts to non-English European languages. Tasks are native-language originals, not translations, authored by professional linguists and domain experts who are native speakers.
Methodology
Each of the 24 EU languages is represented by at least 120 native-authored tasks. Tasks span four genres: legislative text comprehension, administrative correspondence drafting, professional document summarization, and cross-lingual knowledge retrieval. Scoring uses a combination of automated metrics (BLEU, BERTScore) and human evaluation for a representative 15% sample per language.
Get involved
Review the benchmark design, submit a model for the pilot evaluation, or collaborate with our research team.
Contact research →Other benchmarks
The European AI Productivity Index
Assesses whether frontier AI models can perform economically valuable professional tasks in European contexts — EU law, multi-country taxation, industrial standards, and cross-border regulatory analysis.
AI Act Compliance Benchmark
A benchmark evaluating whether AI systems satisfy EU AI Act requirements: risk classification, documentation, transparency, and human oversight at model and system level.