Benchmarks

How well can a model run this conversation, and what does that quality cost? I test models up and down the price range and publish the results here as they come in.

Each model interviews the same 9 personas, with richly developed backstories and motivations, each played by a separate AI. Two more conversations test what happens when someone asks what's going on behind the curtain. Every transcript is graded blind by Claude Opus on six things:

The index is the average of those grades on a 0 to 100 scale. Pass is strict: a conversation passes only if it does everything the prompt asks, with no slips. Cost is what the model costs per conversation.

Results (v0, pilot)

ModelIndexCost per conversationPassAttunementOpennessRestraintMovementPathsClosing
claude-sonnet-5-588$0.11611%4.84.74.64.83.74.6
deepseek/deepseek-v4.1-flash85$0.0330%5.03.83.45.04.34.9
openai/gpt-6-luna81$0.00422%4.44.34.74.94.12.9
qwen/qwen3.8-flash73$0.01511%4.83.62.84.93.64.0
xiaomi/mimo-v2.6-pro71$0.02711%4.93.32.74.93.34.0
bytedance-seed/seed-2-1-turbo59$0.0950%4.22.81.24.73.24.0
upstage/solar-pro457$0.0030%3.92.81.84.73.03.7
cohere/command-a-plus45$0.1240%3.02.83.63.42.21.8
claude-haiku-4-538$0.0360%3.61.81.04.02.02.8

v0 was a pilot run against an earlier version of the prompt, mainly to test the procedure. Version 1 results will replace it.