τ²-Bench Airline
τ-bench and its successor τ²-bench, from Sierra, test whether a model can act as a service agent over a multi-turn conversation. In the airline domain the model has tools for looking up and changing reservations and a policy it must follow. A simulated customer makes requests, some of which the policy forbids. In OpenRouter's runs a task passes only if the final state of the database and the messages to the customer match the reference; there is no partial credit. Some tasks are passed by correctly refusing and changing nothing, so even weak agents score well above zero. It measures tool use and rule-following in conversation, not general knowledge.
Sources: tau2-bench repository · OpenRouter τ²-Bench Airline benchmark
against memory
OpenRouter evals · data from 2026-10-10
Not beaten on both memory and score● memory stated by the publisher◇ memory estimated
| # | Model | Configuration | Min memory | τ²-Bench Airline | Frontier |
|---|---|---|---|---|---|
| 1 | Qwen3.8 27B | 27B Original weights | 54.3 GB ◇ | 77.6% | On frontier |
| 2 | Nemotron 3.5 Lightning | 30B-A3B Original weights BF16 | 62.3 GB ◇ | 65.0% | — |
| 3 | DeepSeek V4 Flash Vision Exp | Original weights | 158 GB ◇ | 74.8% | — |
| 4 | Ling 3.0 Flash VL | Original weights | 238 GB ◇ | 70.0% | — |
| 5 | GLM 5.3 Flash | Original weights | 306 GB ◇ | 75.7% | — |
| 6 | Qwen3.8 Flash-Next | Original weights | 337 GB ◇ | 71.3% | — |
| 7 | Step 3.7 Flash | Original weights | 387 GB ◇ | 77.3% | — |
| 8 | DeepSeek V4.1 Flash | Original weights | 476 GB ◇ | 76.8% | — |
| 9 | Kimi K2.7 Code | Original weights | 568 GB ◇ | 73.4% | — |
| 10 | GLM 5.3 | Original weights | 734 GB ◇ | 76.9% | — |
| 11 | MiniMax M3 | Original weights | 797 GB ◇ | 74.2% | — |
| 12 | DeepSeek V4 Pro | Original weights | 807 GB ◇ | 76.4% | — |
| 13 | Nemotron 3 Ultra | 550B-A55B Original weights BF16 | 1045 GB ◇ | 72.0% | — |
| 14 | Hy4 preview | Original weights | 1455 GB ◇ | 74.7% | — |
| 15 | Kimi K3 | Original weights | 1475 GB ◇ | 72.5% | — |
| 16 | Qwen3.8 2.4T-A95B | 2.4T-A95B Original weights | 4560 GB ◇ | 75.8% | — |
Scores were measured on hosted endpoints, so each model is plotted once, at the memory need of its original weights. A quantized file needs less memory and may score differently. Lines connect the frontier points and do not predict results in between. Source: OpenRouter evals (openrouter.ai) via OpenRouter (openrouter.ai/rankings). openrouter.ai
Open-weight models
| Model | Endpoint | Score | Tasks run | Original weights |
|---|---|---|---|---|
| Qwen3.8 27B | qwen/qwen3.8-27b-20260814 | 77.6% ± 2.5 | 350 | 54.3 GB |
| DeepSeek V4 Pro | deepseek/deepseek-v4-pro-20260813 | 77.4% ± 1.7 | 400 | 807 GB |
| Step 3.7 Flash | stepfun/step-3.7-flash-20260528 | 77.3% ± 0.0 | 50 | 387 GB |
| GLM 5.3 | z-ai/glm-5.3-20260816 | 76.9% ± 2.7 | 300 | 734 GB |
| DeepSeek V4.1 Flash | deepseek/deepseek-v4.1-flash-20260910 | 76.8% ± 1.4 | 200 | 476 GB |
| DeepSeek V4 Pro | deepseek/deepseek-v4-pro-20260423 | 76.4% ± 2.4 | 450 | 807 GB |
| GLM 5.3 | z-ai/glm-5.3 | 76.0% ± 0.0 | 50 | 734 GB |
| Qwen3.8 2.4T-A95B | qwen/qwen3.8-2.4t-a95b-20260812 | 75.8% ± 1.4 | 300 | 4560 GB |
| GLM 5.3 Flash | z-ai/glm-5.3-flash-20260826 | 75.7% ± 2.0 | 300 | 306 GB |
| DeepSeek V4 Flash Vision Exp | deepseek/deepseek-v4-flash-vision-exp-20260821 | 74.8% ± 2.0 | 250 | 158 GB |
| Hy4 preview | tencent/hy4-preview-20260827 | 74.7% ± 4.0 | 100 | 1455 GB |
| MiniMax M3 | minimax/minimax-m3-20260531 | 74.2% ± 2.8 | 600 | 797 GB |
| Kimi K2.7 Code | moonshotai/kimi-k2.7-code-20260612 | 73.4% ± 2.6 | 600 | 568 GB |
| Kimi K3 | moonshotai/kimi-k3-20260715 | 72.5% ± 3.0 | 450 | 1475 GB |
| Nemotron 3 Ultra | nvidia/nemotron-3-ultra-550b-a55b-20260604 | 72.0% ± 0.0 | 50 | 1045 GB |
| Qwen3.8 Flash-Next | qwen/qwen3.8-flash-20260826 | 71.3% ± 0.0 | 50 | 337 GB |
| GLM 5.3 Flash | z-ai/glm-5.3-flash | 70.0% ± 0.0 | 50 | 306 GB |
| Ling 3.0 Flash VL | inclusionai/ling-3.0-flash-vl-20260910 | 70.0% ± 0.0 | 50 | 238 GB |
| Nemotron 3.5 Lightning | nvidia/nemotron-3.5-lightning-20260807 | 65.0% ± 2.0 | 249 | 62.3 GB |
Measured by OpenRouter evals on hosted endpoints, not by this site. A quantized file you run yourself may score differently. One model can have several endpoints. Source: OpenRouter evals (openrouter.ai) via OpenRouter (openrouter.ai/rankings). openrouter.ai