The calibration programme
Every result on this page was measured on synthetic personas driven by a language model. No human being has been measured — by this programme or by any other part of SoulMap. "Detected" means detected in a simulated population whose true trait values we set ourselves.
This page is the run record. What the run showed is on the validation page; how the measurement works is on the methodology page. Here: what was registered, what was executed, what it cost, what broke, and the fences around each artifact.
What "current engine" means, and what was deprecated
On 2026-09-02 every corpus generated before the current engine was deprecated for measurement purposes. Three things had changed at once: the session engine that deals the options, the persona generator (library v5 replaces v4, 3,280 personas each, disjoint identities), and the persona instruction that tells the driver who to be. A figure from an older corpus and a figure from the new one are not comparable, and this programme does not compare them. The deprecated artifacts stay in the repository's history under their original paths; nothing on these pages cites them.
The current instrument is stamped, not assumed. Every row of the confirm
corpus carries the evidence-shelf hash (cfbd7786eb4e), the family-axis
version (v4-facet-axes-2026-08-10 — 27 measurement axes over 29 library
scales; the axis version label is unrelated to the deprecated library
version), the persona-sheet version (v2_dossier) and the realised order in
which that persona read its trait ladder (verify_smoke.py, 6 of 6 checks).
The confirm run
Registration. calibration/phaseA/prereg/WP26_instruction_remediation_confirm.md,
filed before the run. One primary question — do the twelve historically
unreadable families express above chance inside the corpus — with the
statistic, the null construction and the threshold (pooled mean > 0.0822)
fixed in advance. The five classic families as an internal yardstick.
Machiavellianism named for explicit report whatever its sign. Per-family
figures declared descriptive, with the design's detection floor stated
(0.49 median, 0.75 worst, in full-depth units).
Design. 40 personas from the v5 library, stratified by the assignment arm; 50 planned sessions each in chains of 20, so a session starts from the faded state its predecessor left — the engine's continuity machinery sees the persona's own recent past, capped at twenty sessions of history, which is where continuity saturates. Cohorts of eight share dealt scene nodes (one paid generation serves eight personas' decisions, which are their own). One driver model throughout. Sheet order randomised per persona and stamped.
Execution. Batch tier, 1–3 September 2026, 39.6 hours of wall clock across eight process attempts. Two defects were found and fixed during the run, each with the fix proven on the live run before it continued:
- A decorative field in the option-authoring schema (a two-pole "axis" label) was chaining poles indefinitely until the output cap cut the response mid-string; the whole menu then failed to parse. Request drop rate before the fix: 26% (7 of 27 in the measured round); after deleting the field: 0 of 809, then 21 of 5,416 for the remainder (0.4%).
- A sixty-minute batch-job deadline, added to stop a stall from consuming the whole night, raised on expiry and killed a healthy five-hour process at 37%. It now leaves the job's requests to resubmit, as its own comment had always promised, and the deadline was widened to three hours after a 47-minute round cleared the old one with thirteen minutes to spare.
Delivered. 1,831 of 2,000 sessions completed (8.45% attrition against
the registered 15% gate), 3,822 banked choice rows, $123.16
(wave_summary.json, checkpoint_spend.json). The corpus, config, summary
and per-round batch statistics are banked under
calibration/phaseA/wp26/runs/; the registered analysis is
confirm_analysis.py, and its output confirm_result_v1.json records the
rows file's SHA-256 and the git commit it ran from.
169 of 2,000 sessions; registered gate 15% — passes. wave_summary.json
min 42, max 49, none at 50. The fifty-sessions-per-subject milestone misses at 46 — graphq kpi
Never offered: 54, almost all reflection-turn entries. Every family contrasted on ≥ 294 rows — graphq coverage
Against a $0.039 estimate carried into the budget; the difference is the deep-history prompts of chained play
Where the 169 sessions went
Two causes, unlike each other, and both instrumented rather than inferred.
100 sessions — replay drift, all at chain links 2 to 7. A resumed process replays banked decisions and rebuilds each chain link's starting state by re-walking its predecessor. On a chained run that is only sound if the re-walk reproduces the original exactly; across eight restarts it did not always, and the failure signature is unambiguous: every one of these 100 sat at the links that were live across a restart, and died when the banked choice no longer matched a regenerated menu. The mechanism had been seen once before on an earlier-generation run; what was new is that resuming a chained run re-triggers it. The fix — bank the faded hand-off state the moment it is computed and restore it on resume, never re-derive it — shipped on 2026-09-03 with a test that reproduces the drift. It is the reason no persona reached 50 sessions, and the milestone is expected to be reachable on the next run without changing anything else.
69 sessions — engine validation, spread across all links. The model produced a menu the engine refused after its repair attempts (~3.5%, the ordinary floor).
Of the 169, 140 still banked their earlier turns as rows stamped incomplete; 29 produced nothing. Attrition never discards evidence that was already paid for.
What each artifact may and may not support
The dataset census (calibration/DATASETS.md) fences every artifact. For
the current generation:
| Artifact | What it is | May support | May not support |
|---|---|---|---|
confirm corpus (wp26/runs/, 2026-09-03) | 40 v5 personas × ~46 completed sessions, chained, one driver, 3,822 rows | The registered pooled test; descriptive per-family expression and recovery; split-half at full depth; run-integrity and cost figures | Per-family significance claims (detection floor 0.49); any claim about humans; any comparison with a deprecated corpus; separating the trait from how this one driver plays it |
| confirm_result_v1.json | The registered analysis, run once, seeded, hash-stamped | Every number on the validation page | Anything not in it — there is no second analysis |
evidence packs (wp27/packs/) | Per-family dossiers for the twelve weakest families: bands, shelf actions on both sides with offer counts, run figures | The semantic re-verification described on the roadmap | Measurement claims of any kind |
| pre-v5 corpora (r2–r6, wp14, union) | Older engine, v4 library, six drivers | Nothing on these pages — deprecated 2026-09-02 | Any comparison with the current engine |
On the record, and open
- The final holdout is unspent. A 300-persona holdout was frozen by hash on 2026-08-07 and has never been consulted. It belongs to the deprecated library generation and will be re-sealed against v5 before it is spent.
- One driver. The confirm run was driven by a single model family by decision (earlier multi-model corpora had shown no material driver difference on the statistics of interest, but those corpora are now deprecated, so the finding travels as a design choice, not a result). A second driver family is the named reopening condition for the programme's refuted sign-correction lever.
- Per-trait certification needs personas, not sessions. The per-family intervals did not narrow between 27 and 50 sessions; roughly 170 personas would pin each family to ±0.15, and ~380 to ±0.10.
What the run showed is on the validation page; what happens next is on the roadmap.