Open to consulting. Replies in a day.

I build the data foundation your AI product runs on.

Collection, cleaning, evaluation, and the operations around them.

831
samples generated
97.6%
human approved
1
paper, with a DOI
4
pipelines shipped

How I help

Three ways in.

Audit

Find out why your data is hurting the model.

Two weeks, fixed price

Build

A pipeline end to end, collection through evaluation.

Six to twelve weeks

Advise

A second pair of eyes on a team already building.

From four hours a week

What a pipeline looks like.

See it running
  1. Crawl

    Budget capped before the first request.

  2. Chunk

    Streaming, so a large corpus never has to fit in memory.

  3. Generate

    A cost ceiling on every run.

  4. Score

    Every sample rated. 97.6% human approved.

  5. Split

    Grouped by source page first, so no sample can cross the evaluation boundary.

  6. Export

    JSONL, Parquet, or straight to the Hub.

Research

831 samples at 97.6% approved, with splits that cannot leak by construction.
DataForge: A Streaming, Leak-Aware Pipeline for Synthetic LLM Fine-Tuning Datasetsdoi:10.5281/zenodo.22906072The paper and four open questions

Process

Something running in week one.

  1. Look at the data
  2. Fix evaluation first
  3. Ship a thin path
  4. Hand it over

About

I came to data pipelines from the wrong end.

I started by fine-tuning models and kept hitting the same wall: the model was fine, the data was not. Every interesting problem turned out to be upstream.

More about me

Tell me what your data is doing to your model.

Thirty minutes, no preparation needed.

Book a call, opens in a new tab on Google Calendar

Opens in a new tab on calendar.app.google, a Google page.

hello@iantoo.space