Crisis and resilience · Running
A crisis dataset, four generations deep
Four published generations of a supervised fine-tuning set for emergency guidance, each one a versioned artefact rather than a local file.
HuggingFace, Python ·
What it found
Versioning the data, not just the model, is what made two evaluation numbers comparable at all. Without it every improvement claim was unfalsifiable.
- 5datasets published
- 4generations, v1 to v4
The work
What was actually built
Most fine-tuning data lives on somebody's laptop and is described in a README. That is fine until you want to know whether this month's model is better than last month's, at which point you discover the question is unanswerable: the data moved and nobody recorded how.
So every generation of this set is published, in order, under its own name. crisis-companion-sft-v1 came first. crisis-response-training and its v2 followed. crisis-data-v3-sft and v4-sft are the current pair. Four generations, each one fetchable by anybody, including by me in six months when I have forgotten what changed.
That is the whole experiment. Not the content of the data, which is ordinary, but whether treating a dataset as a released artefact rather than an input makes the work legible. It does, and the cost is close to zero.
Still open
What this did not settle.
Does the versioning survive somebody else using it?
Every generation so far was published by the person who made it, who already knew what changed. A collaborator reading only the names would have a thinner story, and nothing here tests that.
See it
Open it, or read the code.
huggingface.co. Opens in a new tab on an external site.