τ²-Bench Airline

τ-bench and its successor τ²-bench, from Sierra, test whether a model can act as a service agent over a multi-turn conversation. In the airline domain the model has tools for looking up and changing reservations and a policy it must follow. A simulated customer makes requests, some of which the policy forbids. In OpenRouter's runs a task passes only if the final state of the database and the messages to the customer match the reference; there is no partial credit. Some tasks are passed by correctly refusing and changing nothing, so even weak agents score well above zero. It measures tool use and rule-following in conversation, not general knowledge.

Sources: tau2-bench repository · OpenRouter τ²-Bench Airline benchmark

against memory

OpenRouter evals · data from 2026-10-10

16 of 16 models
6570758032641282565121024204840968192τ²-Bench Airline (%) ↑Minimum memory (GB) · log scale12345678910111213141516

Not beaten on both memory and score● memory stated by the publisher◇ memory estimated

#ModelConfigurationMin memoryτ²-Bench AirlineFrontier
1Qwen3.8 27B27B Original weights54.3 GB ◇77.6%On frontier
2Nemotron 3.5 Lightning30B-A3B Original weights BF1662.3 GB ◇65.0%—
3DeepSeek V4 Flash Vision ExpOriginal weights158 GB ◇74.8%—
4Ling 3.0 Flash VLOriginal weights238 GB ◇70.0%—
5GLM 5.3 FlashOriginal weights306 GB ◇75.7%—
6Qwen3.8 Flash-NextOriginal weights337 GB ◇71.3%—
7Step 3.7 FlashOriginal weights387 GB ◇77.3%—
8DeepSeek V4.1 FlashOriginal weights476 GB ◇76.8%—
9Kimi K2.7 CodeOriginal weights568 GB ◇73.4%—
10GLM 5.3Original weights734 GB ◇76.9%—
11MiniMax M3Original weights797 GB ◇74.2%—
12DeepSeek V4 ProOriginal weights807 GB ◇76.4%—
13Nemotron 3 Ultra550B-A55B Original weights BF161045 GB ◇72.0%—
14Hy4 previewOriginal weights1455 GB ◇74.7%—
15Kimi K3Original weights1475 GB ◇72.5%—
16Qwen3.8 2.4T-A95B2.4T-A95B Original weights4560 GB ◇75.8%—

Scores were measured on hosted endpoints, so each model is plotted once, at the memory need of its original weights. A quantized file needs less memory and may score differently. Lines connect the frontier points and do not predict results in between. Source: OpenRouter evals (openrouter.ai) via OpenRouter (openrouter.ai/rankings). openrouter.ai

Open-weight models

OpenRouter evals · data from 2026-10-10
ModelEndpointScoreTasks runOriginal weights
Qwen3.8 27Bqwen/qwen3.8-27b-2026081477.6% ± 2.535054.3 GB
DeepSeek V4 Prodeepseek/deepseek-v4-pro-2026081377.4% ± 1.7400807 GB
Step 3.7 Flashstepfun/step-3.7-flash-2026052877.3% ± 0.050387 GB
GLM 5.3z-ai/glm-5.3-2026081676.9% ± 2.7300734 GB
DeepSeek V4.1 Flashdeepseek/deepseek-v4.1-flash-2026091076.8% ± 1.4200476 GB
DeepSeek V4 Prodeepseek/deepseek-v4-pro-2026042376.4% ± 2.4450807 GB
GLM 5.3z-ai/glm-5.376.0% ± 0.050734 GB
Qwen3.8 2.4T-A95Bqwen/qwen3.8-2.4t-a95b-2026081275.8% ± 1.43004560 GB
GLM 5.3 Flashz-ai/glm-5.3-flash-2026082675.7% ± 2.0300306 GB
DeepSeek V4 Flash Vision Expdeepseek/deepseek-v4-flash-vision-exp-2026082174.8% ± 2.0250158 GB
Hy4 previewtencent/hy4-preview-2026082774.7% ± 4.01001455 GB
MiniMax M3minimax/minimax-m3-2026053174.2% ± 2.8600797 GB
Kimi K2.7 Codemoonshotai/kimi-k2.7-code-2026061273.4% ± 2.6600568 GB
Kimi K3moonshotai/kimi-k3-2026071572.5% ± 3.04501475 GB
Nemotron 3 Ultranvidia/nemotron-3-ultra-550b-a55b-2026060472.0% ± 0.0501045 GB
Qwen3.8 Flash-Nextqwen/qwen3.8-flash-2026082671.3% ± 0.050337 GB
GLM 5.3 Flashz-ai/glm-5.3-flash70.0% ± 0.050306 GB
Ling 3.0 Flash VLinclusionai/ling-3.0-flash-vl-2026091070.0% ± 0.050238 GB
Nemotron 3.5 Lightningnvidia/nemotron-3.5-lightning-2026080765.0% ± 2.024962.3 GB

Measured by OpenRouter evals on hosted endpoints, not by this site. A quantized file you run yourself may score differently. One model can have several endpoints. Source: OpenRouter evals (openrouter.ai) via OpenRouter (openrouter.ai/rankings). openrouter.ai