Publication:

The Silk Road to Better Evaluations: Adapting COMET for Low-Resource Uzbek Machine Translation

Loading...
Thumbnail Image

Files

written_final_report.pdf (2.15 MB)

Date

2026-04-27

Journal Title

Journal ISSN

Volume Title

Publisher

Research Projects

Organizational Units

Journal Issue

Access Restrictions

Abstract

Machine translation has reached human-comparable quality for high-resource languages, but low-resource languages remain undeserved by both translation systems and the evaluation metrics used to measure them. This paper examines this gap for Uzbek, a Turkic language with approximately over 35 million speakers, where current MT systems produce output that is not authentically Uzbek. We present a two-stage evaluation framework progressing from word-level to sentence- level translation. At the word level, we curate a 1,000-pair English–Uzbek corpus across 20 topic domains and find that NLLB errors mainly consist of mistranslations and untranslated segments. At the sentence level, we collect 1,592 Direct Assessment ratings from native speakers on a 600-sentence corpus. We evaluate three variants of NLLB and report automatic evaluation metric scores, as well as observations from manual review. Finally, we fine-tune COMET-22 to produce the first learned evaluation metric calibrated for Uzbek (Kendall τ : 0.469 → 0.496, p = 0.017), which shows statistically significant results. We release the corpora, annotations, and fine-tuned model to support further work.

Description

Type of resource

Princeton University Senior Theses

Keywords

Location

Citation