A machine translation evaluation pipeline prioritizing Meta's NLLB-200 model and the FLORES-200 dataset, featuring Kaggle-optimized data loading and UMAP embedding visualization.
-
Updated
May 9, 2026 - Jupyter Notebook
A machine translation evaluation pipeline prioritizing Meta's NLLB-200 model and the FLORES-200 dataset, featuring Kaggle-optimized data loading and UMAP embedding visualization.
What 204 languages pay for the same sentences: the tokenizer tax, measured across 12 tokenizers on FLORES-200 parallel text
Multilingual LLM-as-a-judge framework with anchored pairwise comparison, CI-driven sampling, and human-calibration support.
How many tokens do AI models spend on Burmese? Benchmarks GPT, Claude, Gemini, Llama, Qwen, DeepSeek & more on the same FLORES-200 sentences. Burmese costs up to 11.7× more tokens than English, and more than Thai, Hindi, Vietnamese and Chinese with every tokenizer tested.
From-scratch Transformer NMT for English ↔ Amharic (PyTorch). Data pipeline → BPE → training → beam search → FLORES-200 eval → live demo.
To associate your repository with the flores-200 topic, visit your repo's landing page and select "manage topics."