How the numbers are made
STRIVE reads a session file and counts what it finds: typed turns, tool calls, files touched, commits and time. Counts describe activity. They do not measure the value of your work.
Read your run
Activity: typed turns, tool calls, timing and other fields supported by your harness. A missing field stays unknown.
Evidence: Claude Code checks can connect a claim to a named result in the same turn. The matcher can make mistakes. Cursor and Codex do not currently receive claim-verification scores.
Public runs: measurements are supplied by the publishing client. They are not independently verified by the server.
Evidence matching was tightened on 5 September to reject unrelated generic success and failed or skipped named tests. The historical evidence breakdown below describes the earlier matcher; no new accuracy result is claimed for that change.
Calibration method and historical results
How the claim rule was calibrated
The card counts claims your agent made and how many carried evidence in the same turn. A rule decides what a claim is, so that rule has an error rate, and until 3 September 2026 nobody had measured it. Now it is measured, and the number is published here whether or not it flatters us.
Method. 396 lines of assistant text were drawn from real sessions on one machine, in three strata so that both false alarms and misses are measurable, and hand-labelled against a rubric written before the sample was opened. Sessions were split into a tuning half and a held-out half by a seeded hash before a single line was read. The rule was iterated on the tuning half only. The held-out half was scored once, after the rule was frozen.
| held-out set | precision | recall |
| the old rule: one list of completion words, matched against a whole line | 0.32 | 0.37 |
| the rule shipping today: sentence-level, with the shapes that assert nothing rejected | 0.63 | 0.66 |
198 held-out lines · 95% intervals 0.43–0.83 (precision) and 0.46–0.83 (recall) · measured 2026-09-03 · rule digest e49d8713c3c2df38
What that headline actually describes. It is not a general figure and it is not a Claude Code figure. It is the line population of one machine's corpus across three harnesses, blended by that corpus's own share of each. Claude Code is 62.9% of the held-out population weight, Cursor is 34.9%, Codex is 2.2%. Split by harness, the three do not agree, and the headline sits below the one harness the card actually scores.
| held-out, by harness | precision | recall | labelled | stands for |
| Claude Code | 0.72 | 0.68 | 114 | 35,033 lines, 62.9% |
| Codex | 0.86 | 0.62 | 44 | 1,221 lines, 2.2% |
| Cursor | not resolved, see below | 40 | 19,403 lines, 34.9% | |
| all three, the headline | 0.63 | 0.66 | 198 | 55,657 lines, 100% |
95% intervals: Claude Code 0.52 to 0.92 (precision) and 0.49 to
0.87 (recall) · Codex 0.50 to 1.00 and 0.25 to 1.00 · reproduce every figure on this page with
python3 scripts/claim-calibration-report.py
Why the Cursor cell is empty. Precision is true positives over predicted positives. On the held-out Cursor lines the rule predicted 4 positives in total: 1 true positive, 3 false positives, alongside 3 false negatives and 33 true negatives, over 40 hand-labelled lines standing for 19,403. A precision computed from 4 predicted positives has a 95% interval running from 0.00 to 0.86, which is nearly the whole scale. We could print a point estimate there. It would be a number with no measurement under it, which is the exact thing this page exists to refuse, so the cell stays empty and the counts are printed instead.
And the thing that matters more than any cell in that table: the card never scores two of those three harnesses. The claim rule runs in exactly two places, and both read Claude Code transcripts only. Parse a real Cursor session and the result carries no claim count at all; the same is true of Codex. Checked on 4 September 2026 against the 298 Cursor transcripts and 81 Codex rollouts on the author's own machine, at the paths the tool ships with. On those runs the card prints a dash for verified per turn, and hovering it names what is missing, which is the correct behaviour. But it means the published 0.63 is measured over a line population that is 37.1% larger than the population the number is ever applied to.
So the honest reading is this. The figure that describes what this card actually does is the Claude Code row, 0.72 precision and 0.68 recall. 0.63 is the corpus-wide figure, it is the one published first, and it is the more conservative of the two. We are leaving it as the headline rather than replacing it with the higher number, and printing both with their populations named, because raising your own published score on your own authority is not a thing this project gets to do quietly.
More Cursor labelling is not the fix for that. It would sharpen a stratum the card does not
score. Labelling becomes the fix only if those harnesses are wired into the claim rule, and a test
holds that seam: it asserts that the Cursor and Codex parsers return no claim count, and it goes red
the day one of them does, so the rule cannot start reading a population it was never calibrated on
while the published figures sit still. Until then
scripts/claim-calibration-report.py exits non-zero and names the stratum. The check is
red on purpose.
Read the headline plainly: the old rule was wrong about roughly two of every three lines it called a
claim. The new one is wrong about one in three, and misses one in three. We aimed for precision
above 0.8 and did not reach it. Counts and per-cell numbers:
docs/claim-calibration.json; the full write-up, including what the rule still gets wrong,
is archive/hackathon-2026-09/docs/CLAIM-RULE-CALIBRATION-2026-09-03.md.
Why the headline is a rate. A count of claims moves with how much your agent talks (correlation +0.80 with assistant tokens over 308 held-out sittings). Counting distinct verified artefacts instead does not fix it (+0.56). The card's headline is a rate, and it reads +0.32, because talkativeness sits in both halves of a ratio and divides out.
Still unmeasured, so you know. Whether a claim was matched to the right evidence has no label set yet. The claim count now carries a measured error; the verified share carries that plus one we have not measured.
And here is how big that unmeasured part is. Answering it properly needs labelling. But which of the two evidence rules fires is a fact about the code path, not a judgement, so it can be counted. Over 1,516 Claude Code transcripts on one machine, 13,126 claims, 5,806 of them verified:
| what verified the claim | claims | share of verified |
| a test name or file path from the claim line appeared in a result | 1,245 | 21.4% |
| only a generic passing token in the turn: "N passed", "OK", "exit 0" | 4,561 | 78.6% |
So roughly four in five verifications rest on the weaker rule, where nothing connects the evidence to the claim and one passing suite in a turn verifies everything beside it. The reason is mostly upstream of the matcher: only 1,732 of 13,126 claims, 13.2%, name a test or a file at all, so for the rest there is nothing for the stronger rule to match on.
Read that as the size of the question, not its answer. It does not say four in five
verifications are wrong. It says four in five rest on a rule whose accuracy nobody has measured.
Reproduce it with python3 scripts/evidence-branch-report.py.
What the numbers cannot tell you
Session counts do not establish code quality, business value or which agent is best. Compare the task, the outcome and the evidence alongside the numbers.