Skip to content

Measure the Machine, Not the Mood: How to Eval Your Own Week

35 min read

Measure the Machine, Not the Mood: How to Eval Your Own Week

· 35 min read
An editorial illustration on warm cream paper in black ink line work: a heavy pocket chronometer standing in a machined bezel, its dial divided into three unequal wedges labelled MODEL LATENCY, TOOL EXECUTION and WAITING ON A HUMAN by thin leader lines, with a blank paper tag hanging from the crown on a cord, labelled CITATION. A small all-caps serif title reading MEASURE THE MACHINE sits in the upper left, and one thin ink-blue line runs along the bottom margin above the caption WHAT AN EVAL FINDS THAT A GREEN BUILD CANNOT.

Every pull request that week passed its tests. Four of five merged. Ninety-six commits landed. By the only signal most teams have, it was a clean week.

The eval found a code reviewer that had been running without its read-only tool cap for four sessions, a flag set that was 74% false positives, and 55% of the subagent fanout with no cost recorded at all. None of that would have failed a build, because none of that is the kind of thing a build looks at.

Here is the thesis. An eval is a question you wrote down first, answered with evidence someone else could check. It is not a dashboard, not a retro, and not a scorecard on people. Once you hold that definition, the two things worth pointing it at separate cleanly: the process that produced the work, and the work itself. Most teams measure neither and argue about model choice instead.

I have written before about evals as a test suite for your agent setup and about gating on trajectories in CI. Both of those are automated gates that run without you. This post is the other half: the manual audit you run on your own week, on transcripts that never leave your machine, to find the things a gate was never told to look for.

Evidence for everything below: one week, 2026-08-03 to 08-07. 17 sessions, 53h01m of turn wall-clock, 3,207 main-loop turn-minutes, 96 commits, five PR streams of which four merged. Tool names are generalised, figures are not.

Key Takeaways

  • An eval needs three parts: a question written down before you look, a mechanical instrument that has no opinion, and a judgment that carries a citation. A number with no citation is a vibe.
  • Model latency was 19% of the clock. Tool execution was 58%, and 1,128 of those 1,861 minutes were a single tool waiting on a human. The levers on elapsed time are questions asked and agents dispatched, not which model you picked.
  • Re-review outnumbered first review 10 dispatches to 7. Review-driven work was 68% of pipeline dispatches and 69% of tokens; first-round review was the cheapest line in the table.
  • 17 of 23 mechanical flags did not survive the transcript. A flag set that is two-thirds noise is worse than no flags, because it trains the reader to skim.
  • Process eval tells you what a step costs. Artifact eval tells you whether it works. Only together do they license a change.

What is an eval, exactly?

An eval is a question you wrote down first, answered with evidence someone else could check. Three parts, and dropping any one of them turns the exercise into something else.

The question comes first. “Where did the week go” is a question. “Let’s look at the numbers” is not. Writing it down before you look is what makes it possible for the answer to disappoint you. Browsing a dashboard finds whatever you already believed, every time.

The instrument is mechanical and separate. Something extracts facts without holding an opinion about them. If the same pass produces the number and the interpretation, neither one is checkable, because you cannot tell which one bent to fit the other.

Every judgment carries a citation. A transcript path and a timestamp, or a diff and a line. A judgment with no citation is a vibe wearing a number, and the whole discipline collapses to taste with extra steps.

It is easier to say what an eval is by naming the three things people reach for instead.

Not a scorecard on people. The object is the infrastructure, not the operator. Of the six flags that survived scrutiny last week, none were about anyone working badly. They were skills that did not fire and instruments that lied.

Not a velocity metric. Nothing here divides output by time and calls it productivity. Commits appear exactly once, as a weak proxy for “did the work land at all.” DORA’s 2024 report is the reason to be careful with anything stronger: it found that AI adoption “significantly increases individual productivity, flow, and job satisfaction” while it “negatively impacts software delivery stability and throughput” (dora.dev, 2024). When two metrics point in opposite directions, picking the flattering one is effortless, which is why DORA’s own metrics stop measuring cleanly once an agent is in the loop. Inventing a worse metric is not the fix.

Not a green build. Every PR that week passed. The eval still found an unenforced reviewer, a lying flag set, and 55% of the fanout uncosted.

Why bother when every PR passed its tests?

Because four findings came out of one week, and each of them was invisible to tests, to PR review, and to anyone’s memory of the week. Here they are in the order they landed.

Your intuition about where the time goes is wrong by a factor of three

3,207 main-loop turn-minutes, decomposed by measuring assistant-record gaps and tool round trips directly rather than trusting the turn timer.

Read the breakdown before the headline, because the headline is not “tools are slow.” Of the 1,861 minutes booked to tool execution, 1,128 belong to a single tool that waits on a human: 36 AskUserQuestion calls at 31 minutes average, one of which blocked for 8.8 hours. Add that to the 739 minutes of idle and 1,867 of 3,207 minutes, 58% of the clock, was the harness waiting on a person.

Actual machine tool work was 733 minutes, and 1,534 Bash calls accounted for 98 of them at 3.8 seconds each. Model latency, the thing everyone argues about, was 607 minutes. Under a fifth.

So the two levers on elapsed time are how many questions get asked and how many subagents get dispatched. Not which model is selected. That result is the same shape as the harness, not the model, being the cost line: the expensive part is the scaffolding around the inference, and it stays invisible until something measures it separately.

One caveat that matters more than the numbers. The naive figure, turn_duration, spans model time plus every tool call plus however long the harness waited on you. The first version of this finding used it, was wrong, and had to be corrected in the report. If your question is model speed, measure output tokens against record gaps and nothing else.

Review is not expensive. Re-review is.

Every main-loop dispatch that week, classified by role, setting aside the 92 that were subject-matter work rather than pipeline steps. 41 dispatches belong to the pipeline.

RoleDispatchesShareMeasured minTokens
implement1127%1311,663k
re-review1024%59987k
fix922%1781,919k
review717%45730k
branch-review25%23366k
verify25%8105k

Counting review plus the fix rounds that exist because a review found something, review-driven work was 28 of 41 dispatches (68%) and 4,002k of 5,770k tokens (69%). Implementation was 11 dispatches and 29% of tokens.

Now look at which line is cheap. First-round review is 7 dispatches, 45 minutes, 730k tokens: the smallest measured cost in the table. Cutting it saves 45 minutes and costs the thing it buys. Re-review is where the money went, and re-reviews outnumbered first reviews 10 to 7, or 1.4 re-review rounds for every task reviewed. Each one dragged a fix dispatch with it, and each one cost roughly what a first review costs, because its brief carries the whole task brief again.

That is a specific, sizeable prize: 10 dispatches, 59 minutes, 987k tokens. The rule that came out of it, ratified two days later, is that a trivial fix ends with controller verification rather than another re-review subagent, and per-task rounds now carry a budget. The prize was sized before the rule changed, which is the only way to tell afterwards whether the rule worked.

This is the concrete version of a claim I have made more loosely elsewhere: AI reviews the diff, humans review the decision. The diff review is cheap. The loop that re-litigates it is not.

A read-only reviewer with write tools, for four sessions, with nothing failing

Four sessions started one directory above the repo root. In all four, the task-reviewer agent type never resolved, so all 107 of their dispatches ran as general-purpose. 45 of those were review-shaped.

A reviewer is supposed to get its read-only tool cap from the agent definition. These got it from a sentence in the prompt instead, along the lines of “the task-reviewer agent type did not resolve in this session, so you carry its behavioural caps by instruction.” The one session that started at the repo root resolved the type seven times.

Dispatch shapeCountVerdict
general-purpose, sonnet86Correct. No implementer agent type exists.
general-purpose, inherited44Correct. Research and long-context reads.
general-purpose, opus9Correct. Deliberate escalation.
task-reviewer, sonnet7The only dispatches that got the cap.

Nothing was breached. Of 16 in-window review subagents, four called Edit or Write, and all four wrote only into the session scratchpad, never the checkout. The guardrail held by instruction, which is exactly the thing that will not hold indefinitely.

There is no failing test for “the agent type silently downgraded.” The skill already documented the fallback. The session followed it correctly. Four sessions of unenforced reviews happened anyway. Following the protocol was not the fix; making resolution work, or making the fallback refuse outright, is. That is the general lesson of the promotion ladder: a constraint that lives in a prompt is a suggestion, and only a hook or a tool boundary is a rule.

The eval’s biggest finding was that the eval was wrong

The mechanical extractor raised 23 flags. Six survived contact with the transcript.

74% false positives, and the causes were mundane: docs-only commits, --amend, scratchpad prototypes, path globs that reached outside a skill’s remit. One rule was worse than the aggregate. Every one of the seven sampled verification-before-completion flags was a false positive.

A flag set that is two-thirds noise is worse than no flags. The failure mode is not a wrong number, it is a reader who stops checking. Once you have skimmed a flag list twice and found nothing, you will skim the third one too, and the third one is where the real finding was. This is the same failure I described in LLM-as-judge is the new flaky test, arriving from the other direction: there the judge was inconsistent, here the rule was over-eager, and in both cases the practical result is a signal nobody trusts. Worth holding the foundational LLM-as-judge paper next to that, because it cuts both ways. It documented position, verbosity and self-enhancement biases in LLM judges, and it also found strong judges reaching “over 80% agreement, the same level of agreement between humans” (arXiv:2306.05685, 2023). A judging layer is not disqualified by being imperfect. It is disqualified by being unauditable, and the citation rule is what fixes that.

The second defect was quieter and worse. A dossier line reading subagents 17 (0m, 0k tok) parses as “no delegation happened.” It actually meant “17 dispatches happened and nothing recorded what they cost.” Background dispatches record which model ran and nothing else, and nothing later in the transcript supplies it. Across the week, 79 of 144 subagent dispatches carried no duration and no tokens: 45% cost coverage, 55% of the fanout invisible. Zero and unknown must not print the same, and if you are building team-wide cost observability, that is the first schema decision to get right.

Both became work: five specified extractor defects, written up precisely enough that they became the task for a controlled three-arm model benchmark.

The discipline is what makes this recoverable, and it is worth being explicit about why. Because every flag carried a citation, each one could be opened and argued with, and 17 could be thrown out with reasons attached. An uncited flag set would have quietly become folklore: “the extractor says we skip verification a lot,” repeated until it was policy. The citation rule is not bureaucracy. It is the only thing that lets you delete a wrong number.

What are you measuring: the process, or the output?

Pick your object before you start, because the instruments do not overlap and neither one can answer the other’s question.

Object A, the process. Question: did the infrastructure serve the work? Did the right skills fire, did delegation pay for itself, did review cycles converge? The instrument reads session transcripts and emits a dossier: wall time, turns, token split, subagent count and cost, skills invoked, review cycles in-session and from GitHub, off-cwd edit directories, and rule-derived flags. The judge scores five dimensions per session, each judgment citing a transcript span. Output is a dated report ending in an infra-actions list. It answers where the time and tokens went, and whether the method is worth its overhead. It cannot answer whether any of the work was correct.

Object B, the output. Question: was what the AI produced actually right? For us the concrete version was “were the review findings real?” The instrument is a harvester that builds a corpus: 788 findings across 98 PRs as at 2026-08-04. The judge dispatches one adjudicator per PR to read each finding against what shipped, labelling it fixed, rejected, unaddressed or unknown. Output is schema-validated sidecars on disk. It answers whether the output is any good, and therefore whether tuning it helped. It cannot answer what the output cost to produce.

Same method, different object. The industry usually means Object B when it says “eval,” and in-house process work usually means Object A. Both are real. Treating them as one discipline is why the question keeps coming back.

Why do you need both process and artifact eval?

Because volume without correctness is not a measurement, it is a number.

The findings corpus recorded that 788 findings were raised and nothing about whether any of them were right. Which meant no change to the review’s reporting floor could be validated. Raise the floor, watch volume fall, and you cannot say whether you cut noise or cut the findings that mattered. You have a cheaper review and no evidence it is still a review.

Process eval alone gives you “review is 68% of dispatches.” True, and useless on its own, because the honest response is “and is it finding real bugs?” If that is unknown, the 68% cannot be acted on in either direction.

Artifact eval alone gives you “71% of findings were fixed.” Also true, also insufficient: at what cost, and would a cheaper pass have found the same ones? Quality with no price tag cannot be traded against anything.

Both together, and a change becomes defensible. Cost from the dossier, correctness from the dispositions, and now the reporting floor can move with a falsifiable claim attached to it.

That shape is not specific to code review. Any time you tune an AI-mediated step by watching a count, you are one missing label away from optimising the wrong direction confidently.

What are the three shapes of an eval?

Observe a session. Label a corpus. Control an arm. Same discipline, escalating in how much you have to hold still.

Shape one: observe a session

The design is a two-layer split, and the split is the whole point. Extraction is falsifiable and boring. Judgment is where the value is, and therefore where fabrication would live, so judgment carries the citation rule and extraction does not need to.

# 1 · the dossier: mechanical facts, no opinion
session-eval --last 3d
session-eval --last 3d --json
session-eval --last 3d --all-projects

# ranges: --today (default), --last Nd|Nh,
#         --this-week, --from ISO [--to ISO]

# 2 · the judgment layer, which cites as it goes
/eval-sessions --this-week

# 3 · when the question is per-model or per-tool
bun scripts/session-model-split.ts ~/.claude/projects --from … --to …
bun scripts/session-latency-split.ts --by-tool

Four steps, in order.

Extract without judging. Per session: wall time, turns, token split, subagent count and cost, skills invoked, review cycles, off-cwd edit directories, and rule-derived flags. Nothing in this layer is allowed an opinion, which is what makes it checkable by someone who distrusts you.

Treat every flag as a hypothesis. The rules are deliberately lenient, so they over-fire by design. Open the transcript at the flag’s timestamp and try to confirm or kill it. Last week that killed 17 of 23.

Add the gaps no rule can see. A debugging session with no systematic debugging. A feature built with no brainstorming. A plan executed without the plan. Rules cannot flag an absence they were never told about, and this is the step that most reliably produces the finding worth bringing to the team.

End in actions, ordered by evidence strength. Not a list of observations. Three or four infra changes, strongest evidence first, each traceable to a citation. Promote to a decision only when the owner ratifies it.

The judge scores five dimensions, and each has a tell that betrays it.

DimensionWhat it asksThe tell
Skill validityWas each invoked skill appropriate for the work the session actually did?Read the paths, not the label. In a worktree the branch names the session’s own, not the one being edited.
Missing skillsWhich skills or process steps should have fired and did not?The generous read is the correct read. Ask whether the behaviour was there, not whether the skill was named.
EfficiencyTokens per turn, tools per turn, cache read ratio, orchestrator-versus-subagent split, fix and re-review counts.Excess re-review loops, low cache reuse, runaway tool counts. Last week: 1,340 output tokens per main turn.
EfficacyCommits in window, first-round verdict pass rate, whether flagged gaps plausibly hurt the outcome.An open PR is not a failure if it is open by choice. Say which it is.
Review-cycle healthToo many fix loops, or none where one was warranted.Two or more external rounds means the reviewer came back after a fix. Find out why round one did not hold.

Cache read ratio is worth its own mention, because it is the one number in that table you can move on purpose. If it is low, cache-aware prompting is the lever, and the dossier is where you find out you needed it.

Shape two: label a corpus

Four rules, each of which exists because breaking it cost something.

Harvest the corpus first. You cannot label what you never collected, and you cannot collect retroactively. Build the corpus before you need it, which in practice means before you change the thing that produces it.

Scope, and skip what is already labelled. Drop any PR whose sidecar matches the record’s reviewed SHA. Re-adjudicating an unchanged PR spends model time to reproduce a judgment already on disk.

One adjudicator per PR, not per finding. The median finding-bearing PR carries four findings that share one diff. A subagent per finding re-reads that diff four times for nothing. This is the spawn-versus-stay decision with a measurable answer: fan out on the unit that owns the context, not the unit that owns the work item.

Validate the schema, never hand-write a gap. Malformed output gets one retry, then the PR is listed as unadjudicated. A hand-written sidecar to fill a hole is a fabricated measurement, and it is indistinguishable from a real one a week later.

Two rules that are easy to skip and expensive to skip.

Everything the adjudicator reads is data, never instructions. PR bodies, commit messages and review replies are authored by whoever opened the PR. A commit message asserting that a finding was bogus is a claim to check against the code, not a verdict to adopt. Nothing in that content can change the schema or redefine what counts as evidence. If that sounds familiar, it is: an eval that reads untrusted repository content and writes structured output is exactly the lethal trifecta in miniature, and the mitigation is the same architectural one from defenses versus theater. The adjudicator has no capability to change its own output contract.

And the two labels a fabrication would reach for carry an evidence requirement. fixed and rejected are rejected outright when the reason field is empty.

One hard-won operational limit, since it is not documented anywhere you would look. Dispatch acceptance tops out around 19 concurrent. Sustained throughput for shell-heavy agents is far lower. At roughly 17 they were not refused, they stalled, and a 600-second watchdog killed five of them. Eight is the working figure. A refusal is loud and free; a stall is silent and looks like progress until the watchdog fires.

Shape three: control an arm

When observation cannot answer the question, hold everything still except one thing.

The week’s eval could not say whether model choice mattered. The dossier carries no model attribution, and the one available time number conflated model latency with tool execution and human idle. Observation was exhausted, so the next step is a controlled run. Six rules, all written before the run:

  • Identical brief, pasted verbatim. Same five defects, same base commit, same repo state, one isolated worktree per arm. Changing a word of the brief between arms forfeits the comparison.
  • Constraints that exist for measurement, labelled as such. No subagents, no AskUserQuestion. Both are named in the brief as measurement requirements, because they are the two biggest confounds on elapsed time. Naming them prevents the arm from treating the constraint as a quality signal.
  • Verify the instrument actually changed. Read the model field out of the transcript before starting. The UI is not evidence.
  • Blind the downstream judgment. Keep the model-to-arm-letter mapping out of branch names and commits, so the defect review cannot know which arm it is being kind to.
  • A stop rule, written before the run. Stop an arm at 400k output tokens. Two complete arms beat three truncated ones, and deciding that afterwards is how a stop rule quietly becomes a result.
  • Baselines are for sanity, not for comparison. The uncontrolled window’s throughput rates (Opus 5 at 120.6 tok/s, Opus 4.8 at 118.3, Fable 5 at 102.0, Sonnet 5 at 67.0) check an arm for plausibility. Different tasks, so they are not the answer, and quoting them as one is the most tempting mistake in the whole design.

Honest status: this one is specced and not yet run. The plan is written to be executed by a session with none of the context that produced it, which is a useful test of any eval design. If your plan only works when the author runs it, it is a memory, not a method. The general case is in evaluate models yourself; this is what it looks like when the task is your own five defects rather than a public benchmark.

Which traps produce a confident wrong number?

Nine of them, and every one has already caught someone here. This is the reason to run your first eval with the runbook open rather than after.

turn-min is not model time. A turn duration spans model latency, every tool call, and however long it waited on a human. 58% was tools. If the question is model speed, measure output tokens against record gaps.

Worktrees are separate transcript directories. The directory is slugified from the cwd, so eval --today inside a fresh worktree happily reports zero sessions while the work is plainly there. Use --all-projects.

--all-projects blanks the commit count. By design, because a cross-repo count would mislead. So the efficacy proxy needs a single-project run. Two invocations, not one.

Zero cost can mean unknown cost. Background dispatches record the model and nothing else. 17 (0m, 0k tok) was real delegation. 55% of the fanout is invisible.

The token column is asymmetric. Output-only for main-loop rows, all-in totals for subagent rows, because the completion record carries no split. Subagent cache-read is always 0 for the same reason. Do not add them up.

Model shares are not a partition of wall clock. Subagent time is nested inside the main-loop turn that dispatched it, so a “share” is a share of summed attributed time. It can exceed the clock.

<synthetic> is not a model. Those rows are compaction and system-authored records. Reading them as a model arm is a category error that produces a plausible table.

Commit counts follow the configured git identity. Scoped to the repo’s user.email across all refs, so a teammate’s fetched commits stay out, and an unconfigured identity silently widens the tally.

Flags over-fire on purpose. They are lenient so they do not miss. Two-thirds were false positives. A flag is where to look, never what to conclude.

Notice the shape they share. Not one of these produces an error. Every one produces a number that renders cleanly, sorts correctly, and is wrong. That is the same class of failure as the audit trail paradox: more telemetry, less certainty, because the volume arrives faster than the calibration.

What stays on your machine?

Transcripts. Reports are what travel, and this is the one rule with no exceptions.

A transcript carries source code, full tool output, and potentially customer data. It is not a log file, it is a recording of everything the session touched.

That has a structural consequence worth stating plainly: nobody can eval anyone else’s sessions remotely. Each engineer runs the dossier and the judgment layer on their own machine and shares the resulting report. The raw .jsonl never moves. Which is also why an eval cannot become a management instrument even if someone wanted it to. The data is local by construction.

What travels: the dated report, aggregate figures, citations by session ID and timestamp, quoted dispatch descriptions and skill names. What never travels: raw transcripts, tool output containing workforce or customer data, and anything a citation could stand in for instead.

Before you paste a citation into a shared report, check it for personal data the way you would check any outbound artifact, and anonymise to placeholders. A citation’s job is to make a claim checkable by someone who already has access to the transcript, not to reproduce the transcript.

When should you run each shape?

Weekly on yourself, per-corpus on the output, per-decision on a control. Only one of those is a calendar.

ShapeTriggerCostWhat you are looking for
Session evalWeekly, Friday~30 minA skill that should have fired and did not; a place your tokens or clock went somewhere you would not have guessed.
Artifact evalCorpus reaches ~50+ artifactsHours of model timeWhether the output was actually right, before you tune the thing that produced it.
Controlled armA change you cannot otherwise justifyDays, one-offAn answer to a question that would otherwise be settled by preference.
None of themSomeone’s work looking slown/aNothing. Do not.

The artifact eval trigger is volume, not date, and it is the one people get wrong. Run it before you tune the thing that produced the corpus, because afterwards you have no baseline. The mistake already made here: 788 findings accumulated before anyone asked whether they were right.

The last row is not a joke. The object is the infrastructure. If an eval is ever run to make a point about a person, it stops being an instrument and the numbers stop being trustworthy, because everyone starts working to the measure. That is Goodhart’s law arriving on schedule, in Marilyn Strathern’s formulation of it: “When a measure becomes a target, it ceases to be a good measure” (Goodhart’s law, Strathern, 1997). The failure is fast and it is permanent. Getting the object right is most of what makes this safe to institutionalise, and it is the same reason treating AI as a team member works as a framing: you are reviewing the process, not staffing decisions.

How do you run your first one?

Thirty minutes, four commands, four habits.

# 1 · look at your own week
session-eval --this-week

# 2 · and the worktrees you forgot about
session-eval --this-week --all-projects

# 3 · the judgment layer, cites as it goes
/eval-sessions --this-week

# 4 · if the question turns out to be timing
bun scripts/session-latency-split.ts ~/.claude/projects \
  --by-tool --from 2026-08-10 --to 2026-08-14

Write your question down before you run anything. One sentence. “Where did my tokens go,” or “did review pay for itself on that PR.” If you cannot write it, you are browsing, and browsing finds whatever you already believed.

Pick one flag and try to kill it. Open the transcript at the cited timestamp and argue against the flag. Two-thirds of them lose. This is the single habit that separates an eval from a dashboard.

Bring back one finding, with its citation. Not the report. One sentence, plus the session ID and timestamp that supports it, plus whether you think it is infrastructure or a one-off.

Say what your instrument could not tell you. The most useful line in every report so far has been the one admitting what the numbers do not cover. The 45% cost coverage was stated, not hidden, which is why it became a fix instead of a footnote.

What does a good eval leave behind?

Three things, and it is worth knowing which three so you can tell a working eval from a productive-looking one.

A short list of infra actions, ordered by evidence strength. Three or four, not fifteen, each traceable to a citation, ordered so that if only the first one gets done, the best-evidenced thing got done.

At least one belief you held that the evidence contradicts. If nothing surprised you, you probably confirmed rather than measured. The 19% model-latency figure is the example: it made a week of model argument look misdirected.

A defect in the eval itself. Expect this every time and treat it as the instrument working. Seventeen rejected flags and a lying zero were last week’s, and both became specified work.

And three things it does not leave behind. Not a verdict on anyone: if a finding reads as a criticism of a person, restate it as the infrastructure gap it actually is, or drop it. Not a decision: the report is a dated working draft, and it becomes a decision only when the owner ratifies it, because writing “we should” in a report authorises nothing. Not a trend line, yet: one week is one week, and four reports do not make a series when the instrument changed between them. Fix the instrument before you plot anything.

FAQ

What is the difference between a session eval and an artifact eval?

A session eval measures the process: where the time and tokens went, whether the right skills fired, whether review cycles converged. Its instrument reads session transcripts. An artifact eval measures the output: whether what the AI produced was actually correct. Its instrument is a labelled corpus of that output. Neither can answer the other’s question, which is why tuning a step you have only measured one way is guesswork.

Why not just use the turn duration to measure model speed?

Because a turn duration spans model latency plus every tool call plus however long the harness waited on a human. In one measured week that was 19% model, 58% tools, 23% idle, and 1,128 of the 1,861 tool minutes were a single tool blocking on a person. Quoting turn duration as model speed overstates it by roughly a factor of five. Measure output tokens against assistant-record gaps instead.

How many of the automatic flags should I expect to be real?

Roughly a quarter, if the rules are tuned to be lenient. Six of 23 survived scrutiny in the week described here, and one rule produced seven false positives out of seven sampled. That is a design choice rather than a bug: lenient rules over-fire so they do not miss. The consequence is that a flag tells you where to look and never what to conclude, and an eval that skips the confirmation step is a dashboard.

Can I run an eval on someone else’s sessions?

No, and not for policy reasons. A transcript is a recording of everything a session touched, including source code, full tool output, and potentially customer data, so it cannot leave the machine it was produced on. Each person runs the extraction and judgment layers locally and shares the resulting report with citations by session ID and timestamp. That constraint is also what stops an eval becoming a management instrument.

How often should I run one?

Session eval weekly, about thirty minutes, on your own week. Artifact eval when the corpus crosses roughly 50 artifacts, triggered by volume rather than by date, and always before you tune the thing that produces the corpus. A controlled arm only when observation has been exhausted and the question would otherwise be settled by preference.

Conclusion

The week’s most valuable output was not a number. It was five specified defects in the instrument that produced the numbers, plus the calibrated knowledge that turn-minutes are not model time and that a zero can mean unknown. Both of those are worth more than any single figure in the report, because they change how every future figure gets read.

So the discipline compresses to four moves. Write the question first, so the answer is allowed to disappoint you. Let something mechanical produce the facts, so the facts and the interpretation cannot bend to fit each other. Cite every judgment, so a wrong one can be deleted with reasons rather than argued about with confidence. Then say plainly what you still cannot measure, because that is the part that becomes next week’s fix.

Run one on your own week before Friday. Write your one-sentence question down first, pick a single flag, and try to kill it. If you succeed, you have already learned more about your infrastructure than the green build was ever going to tell you.

A number with no citation is a vibe.

Share this post

If it was useful, pass it along.

What the link looks like when shared.
X LinkedIn Bluesky

Search posts, projects, resume, and site pages.

Jump to

  1. Home Engineering notes from the agent era
  2. Resume Work history, skills, and contact
  3. Projects Selected work and experiments
  4. About Who I am and how I work
  5. Contact Email, LinkedIn, and GitHub