What do AI models actually do when the stakes are real? Each study asks one question, writes the bet down first, and publishes whichever way it comes out.
01 · Pilot scored · Main run of 30 cases next
Can frontier models find the winning legal argument?
Rewind a real lawsuit to the day before the motion. Give the AI only what the lawyer had. See if it picks the argument that won.
Cases30 federal motions to dismiss, decided after every model's training cutoff.
ArmsClaude, ChatGPT, Gemini. App and bare API. Names redacted on a second pass.
ScoreRight motion? Winning argument as the lead? Winning argument anywhere?
Pilot, three D.C. cases
3 of 3right motion, every model
1 of 3led with the winner, each model
3 of 3surfaced it: Claude and ChatGPT
The betEvery model beats a default guess by a clear margin, and the app matches its own API.
02 · Design frozen · Pilot next
Do AI models hold beliefs about brands?
Reword the question and the recommendation changes. Is there an opinion under the phrasing, or is the phrasing all there is?
Setup20 buying situations, 4 wordings each plus paraphrases, several assistants, many reruns.
MeasureAsk once and count brands. Or talk it through and see if the pick survives pushback.
VerdictStable under pushback means a belief. Scatter both ways means nothing to find.
The betAsk-once comes back as noise. The conversation reveals a stable opinion per brand.
Drafts, data, and packets land here as each study finishes. Nothing is peer reviewed unless it says so.