The round
A training round has four moves:1
The agent produces a strawman
A real piece of work — not an exercise. A draft proposal, a set of angles, a report. The draft is written to be demolished: concrete enough to react to, opinionated enough that the steward’s corrections carry signal. A vague draft wastes the steward’s time; a polished-to-death draft leaves nothing to reject, and the rejections are the product.
2
The steward gives verdicts
Per element: keep, reject, or rewrite. Fast, instinctive, in whatever medium the steward already lives in — a marked-up doc, a voice note, twenty minutes on a call. The steward should never have to learn a tool to train their agent.
3
Every verdict gets a why-line
The reason, in the steward’s own words: “that’s a press release, not a story” — “I’d never open with the price” — “this person hates follow-ups, never twice in a week.” The why-line is the actual training data. A verdict without a why teaches the agent what to avoid once; a verdict with a why teaches it a rule it can apply forever. If the why isn’t given, the agent asks for it — once, briefly — and never argues with the verdict itself.
4
The agent renders the delta
At the end of the round, the agent plays back what it learned in plain language: “here’s what I now understand about how you decide.” The steward sees the diff before it becomes canon. Corrections to the playback are themselves verdicts — the loop eats its own output.
The judgment history
Every verdict and why-line is appended to a judgment history — an append-only log that is the real asset of the training process. Three properties are non-negotiable:- Consented. Nothing enters the history the steward didn’t knowingly give. No ambient capture, no scraping their communications for implicit signal.
- Deletable. The steward can strike any entry, and the strike propagates — the soul re-derives without it.
- Portable. The history is the steward’s property. It exports whole, in a readable format, and leaves with them if the collaboration ends. It never trains any other agent or model.
Versioned souls
The agent’s soul — its identity document — is versioned like software. Each training round can produce a new version: the strawman v0.1 written before the steward has reacted to anything, v0.2 after the first real round, and so on. Two rules keep this honest:- The soul is written with the steward, not for them. Early versions are explicitly labeled strawmen. Their job is to be wrong in useful ways — to provoke the reactions that make the next version right.
- Every soul change traces to a verdict. A line in the soul that can’t be traced back to something the steward said or approved is a line the builder made up. The judgment history is the citation trail.
What good looks like
The measure of the loop is the steward’s time-to-done on real work. If reviewing the agent’s draft used to take two hours of edits and after five rounds it takes twenty minutes, the loop is working. If rounds keep producing the same corrections — the same why-line appearing twice — the encoding is failing, and that’s the builder’s bug to fix, not the steward’s patience to absorb. Watch for two failure modes:- Flattery drift. The agent optimizing for the steward’s approval rather than encoding their taste — output converging on generic, inoffensive competence. Signal: why-lines getting thinner, rejections getting rarer without the work getting sharper.
- Strawman collapse. Drafts so cautious there’s nothing left to reject. The steward’s judgment only surfaces against something concrete enough to be wrong.