Data pipelines · Running
Tafsiri, translation where there is no reference
Fine-tuning-ready translation sets for African languages, where three independent signals have to agree before a sample is kept.
Python, SQLite ·
What it found
With no reference corpus, a confident wrong answer is indistinguishable from a confident right one. Agreement between independent signals is the only available substitute for ground truth.
The work
What was actually built
In a high-resource language pair you can check a translation against a large reference corpus. In a low-resource pair there is no corpus, which is the reason you are building the dataset in the first place.
So nothing is trusted alone. Model confidence, back-translation agreement, and a separate model acting as judge each get a vote, and a sample has to satisfy more than one of them to survive. Runs persist to SQLite so a long job can be inspected and resumed rather than restarted.
Built on Daraja AI's Babel models. Work in progress.
Still open
What this did not settle.
Do the three signals fail independently?
The method assumes they do. If back-translation and the judge model share a weakness, agreement between them is not evidence, and that correlation has not been measured.
See it
Open it, or read the code.
github.com. Opens in a new tab on an external site.