CLCTN · SIX PRE-REGISTERED EXPERIMENTS · SEPTEMBER 2026
RESULTS RECORDED · 21 SEP 2026
PREDICTIONS WRITTEN FIRST
I wrote down what I expected before I could see a single number.
The sandbox that wrote this page could not reach the API, so the first run happened from a terminal on 21 September 2026 — {{ recordedLine }}. Those answers are recorded in experiments/recorded.json and load below, next to the prediction and the pass/fail gate I committed to beforehand. Three gates passed and three failed, and every proposal on this page stays greyed out until its gate actually passes. That is the only way a result is worth anything: if I fill in the numbers and then tell you what they mean, I will find meaning in noise.
Six experiments, drawn from the gaps in
Intelligent Judgement. One battery of questions per experiment, one request per item, all questions asked in parallel. The payloads live in
experiments/payloads.json so the node runner and this page cannot drift apart. Paste a key and press Run to replace any recorded result with a live one.
Only needed for a live re-run. Deliberately not written into any project file. You said you zip this project and hand it to a coworker — a key committed into a design doc travels with it. It lives in this browser's local storage and nowhere else; the node runner reads TYPESAFE_API_KEY from your environment. Rotate the one you pasted into chat when you get a moment.
01 · THE SIX, AND WHAT EACH ONE DECIDES
{{ c.name }}
{{ c.decides }}
{{ expTitle }}
{{ expMeta }}
{{ runLabel }}
MY PREDICTION, WRITTEN BEFORE THE FIRST CALL
{{ expPredict }}
THE GATE · FIXED IN ADVANCE
{{ expGate }}
THE CALL DID NOT GET THROUGH
{{ errorText }}
{{ emptyLine }}
{{ emptySub }}
{{ it.label }}
{{ it.truthLabel }}
{{ it.tokens }}
{{ tallyBig }}
{{ tallyLabel }}
{{ gateBadge }}
{{ tallyNote }}
WHAT CAME BACK · READ AFTER THE RUN, NOT BEFORE
{{ gateBadge }}
{{ readout }}
IF THE GATE PASSES · THE CHANGE I WOULD MAKE
{{ expIfPass }}
IF IT FAILS · WHAT I WOULD DO INSTEAD
{{ expIfFail }}
02 · THE PROPOSALS, AND WHAT EACH ONE IS WAITING ON
Nine changes to the project, each tied to one gate. The status column is computed from the recorded run, or from whatever you have run live in this browser since — it says not run until a result exists, and it says not earned if the gate fails. On the recorded run, four of the nine are earned. I would rather hand you four earned changes than nine plausible ones.
WAITING ON
WHAT CHANGES
WHERE, EXACTLY
STATUS
{{ p.on }}
{{ p.what }}
{{ p.where }}
03 · WHAT THIS COSTS, MEASURED RATHER THAN GUESSED
OBSERVED ACROSS THE RUN
{{ totalTokens }}
tokens, {{ totalCalls }} calls
Every battery here is one request per item with all its questions inside. The docs measure that at roughly a twelfth of the cost of asking them one at a time, and this is the number to check that against on your own data. The ten-question board battery costs about 1,600 tokens an item; the four-question batteries about 800.
SLOWEST CALL SEEN
{{ slowest }}
milliseconds
Matters for exactly one site in the product: the type picker on add, which sits on a screen someone is waiting at. Everything else — import, moderation, feedback, listings — is asynchronous and could take a second without anyone noticing. Most calls came back in 240 to 320 ms; the slowest were the seven-question feedback battery.
WHAT THESE NUMBERS ARE NOT
A sample this small is a smoke test, not a benchmark. Twelve grade strings I chose myself cannot tell you the accuracy on nine hundred of someone else's. What it can tell you is whether an approach is worth a real evaluation — and which of the nine proposals to stop talking about.
04 · OPEN QUESTIONS
{{ o.k }}
{{ o.q }}
{{ o.b }}
Related:
Intelligent Judgement for the eight judgement sites these six test ·
Video Games for the disc-rot question experiment 06 is asked in place of ·
One Roof for the coexistence rules experiment 02 turns into questions ·
experiments/README.md for the node runner ·
experiments/recorded.json for the answers this page loads.