Validation

Everything on this page was measured on synthetic personas driven by a language model. No human being has been measured anywhere in this programme. That is the design boundary of the work so far, and it belongs in the first sentence, not a footnote. Whether any result below transfers to people is unestablished, and no amount of further synthetic work can establish it.

This page rests on one run: the pre-registered confirm run of 1–3 September 2026 (calibration/phaseA/wp26/). Every number here comes from its banked analysis, confirm_result_v1.json, computed once on the finished corpus with a fixed seed and stamped with the rows' hash. Earlier generations of the programme — an older corpus engine, an older persona library, and every figure derived from them — were deprecated on 2026-09-02 and are not cited anywhere on this page. By decision, old and new data are not compared: the engine, the persona generator and the persona instruction all changed together, so a "before and after" number would span several confounded changes at once. A one-line note on those earlier runs is at the end.

One rule governs the writing: every number names the artifact it comes from. A number without a named artifact does not exist here.


The run

Personas
40

Drawn from the v5 library (3,280 synthetic personas); each plays 50 planned sessions in chains of 20, so later sessions see the persona's own recent past — confirm_result_v1.json

Sessions completed
1,831 / 2,000

8.45% attrition against the registered 15% gate; 3,822 banked choices — wave_summary.json

Cost
$123.16

$0.067 per session on the batch tier; 39.6 hours wall clock across eight process restarts — checkpoint_spend.json

Driver
one model

gemini-3.7-flash plays every persona; the sheet order is randomised per persona and stamped on every row — wave_config.json

The instrument is the production session engine itself, played by synthetic personas under an isolation harness (replayed model transport, in-memory storage, stubbed images). Each turn deals four options from a versioned evidence shelf whose trait loadings are written by the server, never by the model; the persona's pick is banked as one choice row. The shelf identity (evidence_table_sha cfbd7786eb4e) and the family-axis version (v4-facet-axes-2026-08-10, 27 axes over 29 scales) are stamped on all 3,822 rows (verify_smoke.py, 6 of 6 completion checks passed).

Two boundaries on the corpus itself:

  • Coverage was complete at the family level and thin at the entry level. All 27 families were contrasted on at least 294 rows; 312 of the shelf's 366 entries were offered at least once, and the 54 never offered are almost all reflection-turn entries (graphq coverage). Family-level claims below are on solid footing; per-entry claims are not made.
  • Each persona is bound to one driver, and there is only one driver. No claim on this page separates "the trait" from "how this model plays the trait". A second driver family is a named reopening condition in the programme's refutation record, not a footnote.

The pre-registered question, and its answer

The registration (calibration/phaseA/prereg/WP26_instruction_remediation_confirm.md) asked one question about the twelve families the programme had never been able to read — the Dark Triad, the emotional-intelligence facets, the moral foundations, attachment anxiety, resilience, cognitive flexibility:

Do the twelve express above chance, inside this one corpus?

"Express" is model-free: for each persona and family, how far the chosen option's loading sits above its own menu's mean, on menus where the family was genuinely contrasted; then the correlation of that leaning with the persona's planted trait, across the 40 personas; then the mean over the twelve. The bar — pooled mean above 0.0822 — was fixed before the data existed, as the 95th percentile of a shared-relabelling null.

Result: the twelve express above chance. Pooled mean +0.126. Against this corpus's own null — 5,000 relabellings of which persona each set of choices belongs to, one relabelling shared across all twelve families so their real correlation survives — the null averages +0.0005 with a 95th percentile of +0.078, and the observed value is exceeded by chance in 0.42% of draws (confirm_result_v1.json, H1). The pre-registration allows exactly two statements from this, and this page makes only them:

  1. Under the current stack, the twelve do express above chance.
  2. They sit 0.226 below the five control families — agreeableness, conscientiousness, extraversion, honesty–humility, openness — which pool at +0.352 in the same corpus under the same driver and instruction (confirm_result_v1.json, H2).

No claim of "improvement" is made: the persona instruction and the driver changed together between generations and cannot be separated.

The third registered item, machiavellianism, was named in advance because it had read negative before. It reads negative again: −0.117 in expression and −0.102 in served recovery (confirm_result_v1.json, H3). Its sign chain has been traced by hand and is correct; the negative number is reported, not corrected, and the construct is under the semantic re-verification described on the roadmap.


Family by family

Per-family figures are descriptive, by registration. The design — 40 personas — was sized to answer the pooled question above. Its smallest detectable per-family effect is about 0.49 in full-depth units for the median benchmark family and 0.75 for the hardest, so an individual family that fails to clear zero has not been shown flat; it has not been resolved (confirm_result_v1.json, declared_limits).

Three statistics per family, all against planted truth across the 40 personas. Expression is the model-free leaning above. Recovery is the estimate the product's own estimator produces — in the configuration a user actually gets, every discrimination pinned to 1 and no nuisance terms — with a 95% interval from resampling the personas. Split-half is consistency: score each persona from half their sessions and from the other half (about 23 each), correlate, Spearman–Brown corrected.

familyexpressionrecovery (served)95% intervalsplit-halfconfirmed on its own
honesty–humility+0.727+0.714+0.54 … +0.83+0.86yes
psychopathy+0.600+0.640+0.45 … +0.79+0.13yes
narcissism+0.571+0.520+0.25 … +0.73+0.18yes
self-transcendence+0.524+0.485+0.20 … +0.71+0.45yes
openness to change+0.455+0.424+0.14 … +0.66+0.08yes
authority / subversion+0.387+0.368+0.09 … +0.60+0.46yes
sanctity / degradation+0.378+0.367+0.10 … +0.56+0.04yes
agreeableness+0.344+0.322+0.04 … +0.55+0.19yes
attachment avoidance+0.247+0.317+0.03 … +0.56+0.36yes
stress tolerance+0.317+0.303−0.01 … +0.58+0.40—
extraversion+0.341+0.299+0.04 … +0.53+0.07yes
openness+0.273+0.256−0.05 … +0.54+0.07—
cognitive flexibility+0.238+0.253−0.19 … +0.56+0.57—
loyalty / betrayal+0.235+0.233−0.01 … +0.46+0.01—
empathy+0.180+0.219−0.09 … +0.50−0.42—
care / harm+0.234+0.213−0.13 … +0.50+0.22—
fairness / cheating+0.155+0.178−0.12 … +0.45+0.35—
emotion granularity+0.155+0.121−0.17 … +0.37+0.18—
neuroticism+0.071+0.095−0.29 … +0.46+0.17—
emotional awareness+0.103+0.092−0.21 … +0.38−0.03—
conscientiousness+0.077+0.058−0.31 … +0.36+0.22—
attachment anxiety+0.124−0.009−0.30 … +0.23−0.01—
resilience+0.038−0.017−0.26 … +0.23−0.16—
impulse control−0.098−0.086−0.30 … +0.14+0.15—
machiavellianism−0.117−0.102−0.36 … +0.18−0.15—
flexibility (EI)−0.153−0.185−0.50 … +0.18−0.02—
liberty / oppression−0.284−0.265−0.48 … −0.00−0.37backwards

Source: confirm_result_v1.json, per_family_expression_DESCRIPTIVE and recovery_SERVED_CONFIG. Sorted by served recovery. The same estimator with fitted discriminations from an earlier research fit reaches a median of 0.234 against 0.233 here — the fitted parameters buy nothing yet, which is why none are served (recovery_DESCRIPTIVE).

Read as four groups:

  • Ten families are confirmed individually — the interval sits clear of zero on the positive side. Two of the three Dark Triad scales are among them. Median served recovery across all 27 is 0.233; 21 of 27 read in the planted direction.
  • Thirteen are undetermined: positive or near-zero point estimates, intervals crossing zero. The intervals are about ±0.27 wide, and that width did not narrow between an interim look at 27 sessions and the finish — it is set by the number of personas, not sessions. Resolving these traits one by one is a matter of a larger persona roster, not of engine work.
  • Three read clearly backwards without being resolved: EI flexibility, machiavellianism, impulse control. Their intervals cross zero, but a repeated negative is a blocking finding under the programme's contracts (graph/contracts.yaml, pole_inversion_chain).
  • One is confirmed backwards: liberty/oppression, whose interval sits entirely at or below zero and which did not move between the interim read and the finish.

Those four families are withheld from the product until their language has been re-verified against their scoring keys. The most striking non-result among the undetermined is conscientiousness: a Big Five family, named in every persona's instruction, contrasted on 38% of all rows — the most-offered family after attachment — and still flat at +0.077. That is neither under-mention nor under-exposure; it points at what the scenes' options actually encode for the trait, and it is under the same re-verification.


Reliability at fifty sessions

Population-level recovery and per-person consistency are different questions, and this instrument is much stronger at the first. The split-half column above answers: does the engine describe the same person the same way twice?

value
median split-half reliability, 27 families, served configuration0.149
families at or above 0.307
families at or above 0.502 — honesty–humility 0.86, cognitive flexibility 0.57
families below zero7 — empathy −0.42, liberty/oppression −0.37, resilience −0.16, machiavellianism −0.15, and three within rounding of zero

Source: confirm_result_v1.json, recovery_SERVED_CONFIG.

Stated plainly: at fifty sessions the engine reads a population well on the strong families and reads an individual steadily on only a handful. Two families are not merely uncertain but inconsistent. Every served reading in the product carries its own uncertainty band and a plain confidence word derived from it; the 0.70 bar the programme sets for a decision-grade individual reading is met by one family, honesty–humility.


What went wrong, and what it cost

169 of 2,000 sessions failed. The attrition KPI (≤ 15%) passes at 8.45%; the fifty-sessions-per-subject milestone does not — completed sessions per persona were median 46 (min 42, max 49, none at 50) (graphq kpi).

The cause is instrumented rather than assumed. 100 of the 169 died with the same error, and every one sits at chain links 2–7 — exactly the links that were live across the run's eight process restarts. A resumed process rebuilt each link's starting state by re-walking its predecessor; when that re-walk drifted by a byte, the successor's prompt changed, its cached response missed, the scene regenerated, and the banked choice matched no button. The remaining 69 are ordinary validation failures spread across all links (~3.5%). The fix — banking each link's hand-off state so a resume restores it instead of re-deriving it — shipped on 2026-09-03 with a test that reproduces the drift; the milestone is expected to be reachable on the next run without changing anything else.


What this page does not claim

  • Nothing about human beings. Every artifact above describes model behaviour against synthetic ground truth.
  • Nothing per family beyond description. Only the pooled question was powered; the per-family table is what 40 personas can show, with its detection floor stated.
  • No before/after. Earlier-generation figures are deprecated and not compared against.
  • No decision-grade individual reading except where the table above says so, and the product's serving policy enforces the same rule.

Earlier generations, for the record

Between May and August 2026 the programme ran an earlier corpus engine against an earlier persona library (v4), and published figures from those corpora on this page. Those figures were withdrawn in two rounds — a primary-artifact audit on 2026-08-19, and the deprecation of every pre-v5 corpus on 2026-09-02 once the engine, generator and instruction had all changed. The withdrawn numbers are guarded by a build-time watchlist (graph/claims.yaml) so they cannot quietly reappear on a public surface. The artifacts remain in the repository's history; nothing on this page rests on them.