Technical note · 2026-09-23
DataForge: A Streaming, Leak-Aware Pipeline for Synthetic LLM Fine-Tuning Datasets
831 samples across four websites at 97.6% human-approved quality, with splits that cannot leak by construction.
- 831
- samples generated across four source websites
- 97.6%
- agreement between the scoring rubric and human review
- 4
- source sites, each with a different structure
- 0
- samples crossing the evaluation boundary, by construction
A fine-tuning dataset is only as useful as the evaluation that judges it, and most pipelines quietly undermine their own evaluation at the moment they split the data.
The problem is grouping. If a pipeline generates several samples from one source page and then splits by sample, related samples land on both sides of the train and evaluation boundary. They share phrasing, structure and often the same facts, so the model has effectively seen the answers. The resulting score is inflated, it looks plausible, and nothing in the process flags it. The failure only becomes visible in production, where the model meets text it genuinely has not seen.
This note describes a pipeline built so that cannot happen. Samples are grouped by source page before anything is split, and the groups are ordered by a stable hash rather than shuffled with a seed. That second detail matters more than it looks: a hash means the same corpus produces the same split on every machine and every rerun, so two evaluation numbers are actually comparable. Seeded shuffling appears to do the same thing and stops being reproducible the moment somebody changes language version.
The pipeline streams rather than loading a corpus into memory, so the ceiling on input size is disk rather than RAM, and every run carries a cost ceiling set before the first request. Each generated sample is scored, and the score is checked against human judgement rather than assumed.
Across four websites it produced 831 samples at 97.6% human-approved quality. That number is a measurement of agreement between the scoring rubric and a person reading the samples, not a model evaluating itself.
The note also covers where the approach stops working, which is the part most write-ups leave out and the part worth reading.
Open questions
What it did not settle.
Each of these is a limitation the paper states plainly. They are here because they are what the next piece of work is about, and because a method that only ever works is a method nobody has tested.
How thin can a source be before synthetic generation stops helping?
The pipeline assumes a page carries enough substance to generate from. On thin content it still produces samples, and those samples score well while teaching the model very little. Scoring catches bad samples, not shallow ones, and there is no measure here that separates the two.
What replaces a paywalled source without breaking its terms?
A great deal of the best domain writing sits behind a paywall. Crawling it is not an option. The open question is whether an openly licensed substitute produces a dataset that transfers, or whether the result quietly describes a different domain than the one intended.
Where is the ceiling on synthetic data?
Quality held at 97.6% across 831 samples from four sites. That says nothing about forty sites, or about the point at which generated samples start reinforcing each other rather than adding information. Finding that ceiling needs a measurement this work does not yet have.
Does group-aware splitting change the numbers people report?
Splitting by source page rather than by sample is the correct thing to do, and it is not what most pipelines do. The size of the difference is the interesting part. If inflation from naive splitting is small, this matters less than claimed here. If it is large, a good deal of published evaluation is optimistic.
The source
Read it, or run it.
doi.org and github.com. Both open in a new tab on an external site.