Research
Published, with the code and the open questions.
What exists, what it measured, and what it did not settle. The last of those is the honest measure of where the work is.
Published work
- PreprintPreview2026-10-03VAPF: A Vendor-Neutral Specification for Real-Time Conversational Agent PipelinesAbstractA contract-based specification unifying transport, speech, tool calling, memory and adaptive safety behind versioned interfaces.v0.4.1-draftRead the findings
- Technical notePreview2026-09-28Vision Dataset Labeler: A Requirements Specification for Human-First VLM DatasetsAbstractA labelling tool for vision-language fine-tuning where AI output is always a draft, spend is bounded before it happens, and no work is ever lost.v0.1 draftRead the findings
- Technical note2026-09-23DataForge: A Streaming, Leak-Aware Pipeline for Synthetic LLM Fine-Tuning DatasetsAbstract831 samples across four websites at 97.6% human-approved quality, with splits that cannot leak by construction.doi:10.5281/zenodo.22906072Read the findings
Research areas
Four areas, one at a time.
Narrow on purpose. A question taken to a measurable answer is worth more than four taken to an opinion.
Model and data evaluationCurrent
Whether a number can be trusted, and what it costs when it cannot. The published paper and its code sit here.
Physical AI
Models that act on the world rather than describe it, where a wrong answer moves something.
Robotics
Hardware in the loop, so the same control code runs against a simulator and against the real thing without being rewritten.
Aviation
Where the tolerance for an unexplained failure is lowest, which makes it the hardest honest test of everything above.
A team of two, one question at a time.
We finish before we start something else. That is slower than it sounds, and it is the reason anything here carries a measurement rather than a claim.