Technical note · 2026-09-28

Vision Dataset Labeler: A Requirements Specification for Human-First VLM Datasets

A labelling tool for vision-language fine-tuning where AI output is always a draft, spend is bounded before it happens, and no work is ever lost.

3
principles the design answers to
12
functional requirement areas

Fine-tuning a vision-language model needs image and text pairs whose text is accurate and consistent. The established annotation tools aim at boxes and class labels, and modern vision-language models learn from free-form image and text supervision instead, so the useful work is the text and the human review that makes it trustworthy.

This specifies that tool. A project holds images, an instruction and its labels. Labelling is human by default: a new project has AI switched off and nothing reaches any model until somebody opts in. Assistance from a configurable teacher model is optional, per image or as a cost-controlled batch.

Three principles do most of the work. Human first, so AI is off until asked for. Drafts are not truth, so only human-approved labels are exported. Originals are sacred, so source images are never modified and edits are stored as metadata applied at export.

Spend is explicit and bounded, with an estimate before a batch and a ledger after it. Every action is persisted and every long job is resumable, because the failure that loses an afternoon of careful labelling is the one that stops a tool being used.

It is a requirements document, at v0.1 draft, with its own open questions section unresolved.

Open questions

What it did not settle.

Each of these is a limitation the paper states plainly. They are here because they are what the next piece of work is about, and because a method that only ever works is a method nobody has tested.

  1. Does bounded cost survive a real batch?

    Estimation before a run and a ledger after it is the design. Whether the estimate is close enough to be trusted, and what a person does when it is not, is the thing only a real run against real pricing answers.

  2. Is human-by-default sustainable at volume?

    The principle is that only approved labels are exported. It is easy to hold at a hundred images and the pressure to batch-approve grows with the corpus, which is exactly when the guarantee matters most.

The source

Read it, or run it.