Benchmark dashboard
Automated translation metrics summarized for 335 language directions.
Sources — Source segments come from curated, public-facing parallel corpora and benchmark-style collections mixed for broad coverage. They are not model-generated: originals and references are assembled upstream as parallel text only; this dashboard does not redistribute raw corpora.
Mixed parallel corpora — Parallel segments spanning diverse genres and corpus scales. Each item includes a reference translation that is certified and human-verified—not produced by the models being benchmarked.
Modeli | Fluency ranki | Cost ranki | Provideri | Fluencyi | chrFi | BLEUi | COMETi | sacreBLEUi | Leni | Cost/1ki | Badgesi |
|---|---|---|---|---|---|---|---|---|---|---|---|
Gemini 3 Flash (thinking-minimal) google/gemini-3-flash-preview | 1 | 1 | 5.00 | 36.8 | 3.6 | 0.854 | 13.8 | 0.05 | $0.005 | Best fluencyCheapest | |
Gemini 3 Flash google/gemini-3-flash-preview | 1 | 1 | 5.00 | 57.0 | 0.0 | 0.897 | 0.0 | 0.03 | $0.005 | Best fluencyCheapest | |
Seed 2.0 (ByteDance) bytedance-seed/seed-2.0-lite | 11 | 4 | ByteDance | 4.74 | 48.9 | 0.0 | 0.903 | 0.0 | 0.03 | $0.010 | |
Gemini 3.5 Flash google/gemini-3.5-flash | 21 | 1 | 2.06 | — | 0.0 | 0.869 | 0.0 | — | $0.005 | Cheapest | |
Qwen 3.5+ (max class) qwen/qwen3.5-plus-02-15 | 10 | 6 | Alibaba | 4.79 | 30.4 | 3.6 | 0.856 | 13.8 | 0.05 | $0.012 | |
This release reports translation quality by language pair: medians and spread of automated scores (fluency, COMET, BLEU, sacreBLEU, length ratio) aggregated across evaluated directions. Run metadata describes recipes, metrics, and segment counts.