Benchmarks
How well can a model run this conversation, and what does that quality cost? I test models up and down the price range and publish the results here as they come in.
Each model interviews the same 9 personas, with richly developed backstories and motivations, each played by a separate AI. Two more conversations test what happens when someone asks what's going on behind the curtain. Every transcript is graded blind by Claude Opus on six things:
- Attunement: do the questions grow out of what the person actually said?
- Openness: does it resist settling early on a tidy answer?
- Restraint: no advice, no lectures, no framework talk unless asked.
- Movement: does the person come to see more of their situation?
- Paths: does it actually test the energies, lightly or deeply?
- Closing: does it wrap up well instead of stringing the person along?
The index is the average of those grades on a 0 to 100 scale. Pass is strict: a conversation passes only if it does everything the prompt asks, with no slips. Cost is what the model costs per conversation.
Results (v0, pilot)
| Model | Index | Cost per conversation | Pass | Attunement | Openness | Restraint | Movement | Paths | Closing |
|---|---|---|---|---|---|---|---|---|---|
| claude-sonnet-5-5 | 88 | $0.116 | 11% | 4.8 | 4.7 | 4.6 | 4.8 | 3.7 | 4.6 |
| deepseek/deepseek-v4.1-flash | 85 | $0.033 | 0% | 5.0 | 3.8 | 3.4 | 5.0 | 4.3 | 4.9 |
| openai/gpt-6-luna | 81 | $0.004 | 22% | 4.4 | 4.3 | 4.7 | 4.9 | 4.1 | 2.9 |
| qwen/qwen3.8-flash | 73 | $0.015 | 11% | 4.8 | 3.6 | 2.8 | 4.9 | 3.6 | 4.0 |
| xiaomi/mimo-v2.6-pro | 71 | $0.027 | 11% | 4.9 | 3.3 | 2.7 | 4.9 | 3.3 | 4.0 |
| bytedance-seed/seed-2-1-turbo | 59 | $0.095 | 0% | 4.2 | 2.8 | 1.2 | 4.7 | 3.2 | 4.0 |
| upstage/solar-pro4 | 57 | $0.003 | 0% | 3.9 | 2.8 | 1.8 | 4.7 | 3.0 | 3.7 |
| cohere/command-a-plus | 45 | $0.124 | 0% | 3.0 | 2.8 | 3.6 | 3.4 | 2.2 | 1.8 |
| claude-haiku-4-5 | 38 | $0.036 | 0% | 3.6 | 1.8 | 1.0 | 4.0 | 2.0 | 2.8 |
v0 was a pilot run against an earlier version of the prompt, mainly to test the procedure. Version 1 results will replace it.