Data pipelines · Running

Tafsiri, translation where there is no reference

Fine-tuning-ready translation sets for African languages, where three independent signals have to agree before a sample is kept.

Python, SQLite ·

What it found

With no reference corpus, a confident wrong answer is indistinguishable from a confident right one. Agreement between independent signals is the only available substitute for ground truth.

The work

What was actually built

In a high-resource language pair you can check a translation against a large reference corpus. In a low-resource pair there is no corpus, which is the reason you are building the dataset in the first place.

So nothing is trusted alone. Model confidence, back-translation agreement, and a separate model acting as judge each get a vote, and a sample has to satisfy more than one of them to survive. Runs persist to SQLite so a long job can be inspected and resumed rather than restarted.

Built on Daraja AI's Babel models. Work in progress.

Still open

What this did not settle.

  1. Do the three signals fail independently?

    The method assumes they do. If back-translation and the judge model share a weakness, agreement between them is not evidence, and that correlation has not been measured.

See it

Open it, or read the code.