Data pipelines · Running

DataForge, splits that cannot leak

A streaming, leak-aware pipeline that groups by source before it splits, so an evaluation score means what it claims.

Python ·

What it found

Grouping before splitting is the whole difference between a score and a guess. The failure is invisible in testing and only shows up in production, on text the model genuinely has not seen.

  • 831samples generated
  • 97.6%human approved

The work

What was actually built

A pipeline that generates several samples from one source page and then splits by sample has already broken its own evaluation. Related samples land on both sides of the boundary, they share phrasing and facts, and the model has effectively seen the answers. The score looks plausible and is wrong.

DataForge groups by source page before anything is split, and orders the groups by a stable hash rather than a seeded shuffle. The second detail matters more than it looks: a hash gives the same split on every machine and every rerun, so two evaluation numbers are actually comparable.

This one produced a paper, which carries the method and the limitations properly.

See it

Open it, or read the code.