Methodology
Everything on this page was measured on synthetic personas driven by a language model. No human being has been measured. That is the programme's design boundary, stated first because every other claim depends on it. We built an instrument that can be scored against known ground truth — synthetic personas whose generating trait values we hold exactly — so that its failures would be measurable before any human study. Where a limit exists, this page reports it with the same precision as the successes.
One rule governs the writing: every number names the artifact it comes
from. A number without a named artifact does not appear here. All numbers
on this page come from one run — the pre-registered confirm run of
September 2026 (calibration/phaseA/wp26/confirm_result_v1.json). Earlier
corpora, generated by an earlier engine over an earlier persona library,
were deprecated on 2026-09-02 and are not cited.
One estimator, served and measured
SoulMap has one measurement path, and an earlier version of this page had to explain why the served path and the measured path differed. Since 2026-08-27 they do not: the estimator that produces a user's trait scores is the estimator the calibration programme measures.
It runs in a deliberate uncalibrated configuration: every family's
discrimination is pinned to 1 and no nuisance terms are fitted, because no
calibration manifest is mounted in production. That is a measured choice,
not a placeholder. On the confirm corpus the served configuration recovers
planted traits at a median correlation of 0.233 across 27 families; the
same estimator with fitted discriminations and nuisance terms from an
earlier research fit reaches 0.234 — the same number to within noise
(confirm_result_v1.json, recovery_SERVED_CONFIG and
recovery_DESCRIPTIVE). Fitted parameters buy nothing yet, so none are
served. Everything below describes the instrument in the configuration a
user actually gets.
The shelf: the server writes the psychometrics, the model writes the prose
The psychometric core of a session is not the language model. It is a
versioned evidence shelf of 366 archetype entries — content hash
7bda5d9bf355 since the wording revision of 3 September 2026, stamped on
every session record and every banked corpus row (the confirm corpus below
was banked under its predecessor, cfbd7786eb4e; the revision re-worded
four option texts and changed no trait loading. A staging cue added to two
further entries in the same revision was withdrawn the same day, after a
registered re-measurement found it cut how often that pair was offered by
about three quarters). Each entry carries authored measurement metadata: signed trait
loadings per family, an intent tag, a difficulty rating and a polarity;
registered opposite-pole pairs carry a partner.
At each turn the server selects the four answer options by a deterministic, seeded search over the shelf under hard constraints — distinct primary constructs per menu, a cap per registry family on action turns, an opposite-pole gate for the registered pairs — with soft rules that relax down a fixed ladder rather than fail the turn. The language model then writes the scene and the wording of the four buttons; it never assigns psychometrics. After rendering, the server re-attaches each button's trait loadings, intent tag and difficulty from the table. Trait identity on every button is correct by construction, never model-authored.
The model writes the prose. The server writes the psychometrics.Two honest consequences of this design:
- The design matrix is the shelf's own declaration of what each button means. A recovery figure is therefore a joint test of the driver and the label registry — not of the driver alone. Where a family reads backwards, the first suspect is the language of its options against its scoring key, and the programme's contract treats a repeated negative as blocking until that has been traced by reading, not by statistics.
- Delivery counts, not shelf counts. On the confirm corpus, 312 of the 366 entries were offered at least once; the 54 never offered are almost all reflection-turn entries. Every one of the 27 families was genuinely contrasted — offered with differing loadings on the same menu — on at least 294 of 3,822 rows, and on as many as 1,908 (attachment avoidance).
The choice model
The unit of analysis is one dealt menu: four shelf-backed options plus a free-text alternative, and the observed pick. The measurement model is a conditional logit. The utility of option j for person i is:
U(i, j) = Σ over families f of λ_f · s(j, f) · θ(i, f)
where s(j, f) is the shelf's signed loading of option j on family f and θ(i, f) is the person's trait score on that family. In the served configuration every λ_f = 1 and there are no nuisance terms; choice probabilities follow by softmax over the alternatives. A person's θ is the maximum-a-posteriori estimate under a standard-normal prior given their banked choices, with a posterior standard error per family. That standard error is what the product turns into the interval and the confidence word beside every reading.
Two things the served model deliberately does not do, and why:
- No fitted discrimination. Measured to add nothing on the current engine (0.233 versus 0.234 above); on earlier corpora it measured slightly worse. Pinning λ to 1 also removes a dependency on any fit artifact.
- No position or difficulty terms. Real position bias exists in synthetic choice; the served model absorbs none of it. This is a known simplification, and the in-run comparison above says it costs nothing measurable at present.
The synthetic-persona programme
Ground truth. Personas come from a library sampled by a Gaussian-copula
forge with no language model involved: 3,280 personas in the v5 library
(persona_library_v5.jsonl plus a 400-persona mid-grid supplement), each
with 29 generating trait scales that project onto 27 measurement axes — the
only two-to-one collapses are the two Schwartz bipolar contrasts
(family_map_v4.json, the axis map's version label; unrelated to the
deprecated library v4). A persona's θ on a family is known exactly, which
is the entire point of synthetic ground truth: recovery is scored against
truth, not against another questionnaire. The trait dimensions are named
after public instruments; no instrument was administered to anyone.
The persona instruction. A deterministic sheet of banded, categorical
trait renderings — never the raw trait vector. The v2 sheet names every
construct (an earlier sheet named only the top two moral foundations and one
attachment word, and its under-mention was measured), reads the trait
ladder in a randomised order per persona that is stamped on every row, and
had every band's wording placed by blind readers at the level it claims —
fourteen of fifteen upper emotional-intelligence bands place exactly
(calibration/phaseA/wp24/). Age and gender condition the persona sampling
but are never conveyed to the driver.
The driver. One model family plays every persona on the confirm run. Earlier multi-driver corpora are deprecated, so nothing on this page separates "the trait" from "how this model plays it"; a second driver family is the named reopening condition for the programme's refuted sign-correction lever.
Run integrity. 40 personas × 50 planned sessions in chains of 20, so a
session begins from the faded state its predecessor left. 1,831 of 2,000
completed (8.45% attrition against a 15% gate; wave_summary.json);
completed sessions per persona median 46. The failure that stopped every
persona short of 50 was located to the resume path of chained runs,
instrumented, fixed and tested (the calibration
page has the account).
The testing discipline
Pre-registration. The confirm run's primary question, statistic, null
construction and threshold were filed before the run
(calibration/phaseA/prereg/WP26_instruction_remediation_confirm.md), and
the analysis was run once on the finished corpus with a fixed seed. Two
earlier registered tests in this programme missed their thresholds and were
published as misses; that is what the discipline is for.
The model-free statistic. To test whether a family expresses — whether the persona's choices lean the planted way at all, before any estimator is involved — the programme uses expression: for each persona and family, how far the chosen option's loading sits above its own menu's mean, on menus where the family was genuinely contrasted, correlated with the planted trait across personas. It touches no fitted parameter, so an estimator cannot flatter it.
The shuffled null. "Could this be chance?" is answered by keeping every scene, option and choice exactly as it was and relabelling only which persona each set of choices belongs to — 5,000 times, one relabelling shared across all twelve families so their real between-family correlation survives. Relabelling each family separately would produce a fictitiously tight bar the corpus clears by construction.
Bootstrap intervals per family. Recovery per family carries a 95% interval from resampling the 40 personas. At this roster the intervals are about ±0.27 wide, and they did not narrow between an interim read at 27 sessions and the finish: their width is set by the number of personas, so per-trait certification is a persona-count question (~170 personas for ±0.15, ~380 for ±0.10), not an engine question.
Split-half consistency. Score each persona from half their sessions and from the other half, correlate the two, and apply the Spearman–Brown correction: does the engine describe the same person the same way twice? Only meaningful at depth — at 27 sessions the halves are too thin — which is why the run was not cut short of 50.
Refutations, on the record. Levers the programme has tested and found
dead are recorded with the evidence and a reopening condition
(graph/refuted.yaml): adaptive item selection, item purity, item
retirement, loading sharpening, statistical sign correction, and content
re-authoring for families the persona does not express. Three of them
looked large in-sample and reversed sign held out; every gain is
cross-validated before it is believed.
What the discipline measured
Expression of the twelve benchmark families against a bar of 0.082 fixed in advance; p = 0.0042 against the run's own shuffled null — confirm_result_v1.json
Median across 27 families; 21 of 27 read in the planted direction; 10 confirmed individually; 1 confirmed backwards and withheld
Median per-person consistency at ~46 sessions; honesty–humility 0.86 is the one family above the 0.70 bar
Smallest per-family effect the 40-persona design can resolve (0.75 for the hardest family); per-family figures are descriptive by registration
Detection and measurement are different bars, and this instrument is much stronger at the first. Reading a population is well established on the strong families; reading an individual steadily is established on one. The validation page carries the full per-family table.
What this methodology does not claim
- No human data exists anywhere in the programme. Every θ, every interval, every reliability figure was computed on synthetic personas.
- No per-family significance. The design was powered for one pooled question; the per-family table is description with its detection floor stated.
- No before/after. The current engine is not compared with the deprecated corpora, because the engine, the persona generator and the instruction changed together.
- No conformal coverage. No coverage guarantee has been measured for the served intervals; they are posterior intervals under a standard-normal prior and are labelled as such.
- No measurement invariance has been tested. It cannot be, before a real-user cohort exists.
- The persona population's Σ-fidelity check missed its registered bar — relative Frobenius distance ≈0.12 against a 0.10 tolerance — and the measurement stands as a miss.
The calibration page carries the run record; validation carries the results and the withdrawn-figure watchlist.