- Developed a first-of-its-kind evaluation framework for Indic Large Language Models, addressing the current English-centric bias in performance measurement across tasks like machine translation and cross-lingual question answering.
- Theframework, compatible with various datasets, integrates TGI to speed up evaluation by 45%, offering metrics like accuracy,
F1 score, perplexity, ROUGE, and BLEU. It also supports multi-cloud environments for easy execution
The Indic LLM Leaderboard utilizes the indic_eval evaluation framework , incorporating SOTA translated benchmarks like ARC, Hellaswag, MMLU, among others. Supporting 7 Indic languages, it offers a comprehensive platform for assessing model performance and comparing results within the Indic language modeling landscape.