Is more thinking worth the tokens?

Most current models let you dial how long they reason before answering. More effort usually means more output tokens, more waiting and a bigger bill. These charts show what it buys on Salesforce work.

Vendors name effort differently (low, medium, high, xhigh, thinking on or off). Forcebench records the vendor's own setting and maps it to a common tier (off, low, medium, high, max) so models can be compared. Each line is one model at one quantisation on one engine.

Every answer has a 32,768-token output budget, reasoning included. A config that thinks past it without answering fails the task, so overthinking shows up in the score. The No answer column shows how often that happened.

Score by effort tier
Overall pass@1; vertical bars are 95% confidence intervals.
Full benchmark: pass@1 by reasoning effort tierQwen3.8 27B · AWQ-INT4: low 19%, medium 20%, max 19%0%25%50%75%100%LowMediumMaxReasoning effort tierpass@1Qwen3.8 27B · AWQ-INT4
Show the numbers
ModelEffort (tier)pass@195% CITasksNo answerOutput tokensReasoning tokensLatency
Qwen3.8 27B · AWQ-INT4low (low)19%15–23272<1%3.7k3.1k27 s
Qwen3.8 27B · AWQ-INT4medium (medium)20%17–242722%4.7k3.6k35 s
Qwen3.8 27B · AWQ-INT4xhigh (max)19%14–2327225%19k10k149 s
Score against output tokens
Each point is one effort setting. Further right means more tokens per task, which means more cost and latency.
Full benchmark: pass@1 against mean output tokens per taskQwen3.8 27B · AWQ-INT4: 3.7k tokens 19%, 4.7k tokens 20%, 19k tokens 19%0%25%50%75%100%5k10k20kMean output tokens per task (log scale)pass@1lowmediumxhighQwen3.8 27B · AWQ-INT4
Show the numbers
ModelEffort (tier)pass@195% CITasksNo answerOutput tokensReasoning tokensLatency
Qwen3.8 27B · AWQ-INT4low (low)19%15–23272<1%3.7k3.1k27 s
Qwen3.8 27B · AWQ-INT4medium (medium)20%17–242722%4.7k3.6k35 s
Qwen3.8 27B · AWQ-INT4xhigh (max)19%14–2327225%19k10k149 s
Qwen3.8 27B · AWQ-INT4

low → xhigh: −1 pt

From 19% to 19% for 5.0× the output tokens and 5.6× the latency. The intervals overlap, so the difference is within noise. At xhigh, 25% of answers never arrived: the model used its whole budget thinking.

Want this for your team?

Leo helps Salesforce teams run AI they own: open-weight models on infrastructure you control, tested on your kind of work before you rely on them. A private benchmark of your shortlist is a fixed fee.

Get a private benchmark

Or email leo@azl.au · azl.au