Study · Reasoning effort
Is more thinking worth the tokens?
Most current models let you dial how long they reason before answering. More effort usually means more output tokens, more waiting and a bigger bill. These charts show what it buys on Salesforce work.
Vendors name effort differently (low, medium, high, xhigh, thinking on or off). Forcebench records the vendor's own setting and maps it to a common tier (off, low, medium, high, max) so models can be compared. Each line is one model at one quantisation on one engine.
Every answer has a 32,768-token output budget, reasoning included. A config that thinks past it without answering fails the task, so overthinking shows up in the score. The No answer column shows how often that happened.
Show the numbers
| Model | Effort (tier) | pass@1 | 95% CI | Tasks | No answer | Output tokens | Reasoning tokens | Latency |
|---|---|---|---|---|---|---|---|---|
| Qwen3.8 27B · AWQ-INT4 | low (low) | 19% | 15–23 | 272 | <1% | 3.7k | 3.1k | 27 s |
| Qwen3.8 27B · AWQ-INT4 | medium (medium) | 20% | 17–24 | 272 | 2% | 4.7k | 3.6k | 35 s |
| Qwen3.8 27B · AWQ-INT4 | xhigh (max) | 19% | 14–23 | 272 | 25% | 19k | 10k | 149 s |
Show the numbers
| Model | Effort (tier) | pass@1 | 95% CI | Tasks | No answer | Output tokens | Reasoning tokens | Latency |
|---|---|---|---|---|---|---|---|---|
| Qwen3.8 27B · AWQ-INT4 | low (low) | 19% | 15–23 | 272 | <1% | 3.7k | 3.1k | 27 s |
| Qwen3.8 27B · AWQ-INT4 | medium (medium) | 20% | 17–24 | 272 | 2% | 4.7k | 3.6k | 35 s |
| Qwen3.8 27B · AWQ-INT4 | xhigh (max) | 19% | 14–23 | 272 | 25% | 19k | 10k | 149 s |
low → xhigh: −1 pt
From 19% to 19% for 5.0× the output tokens and 5.6× the latency. The intervals overlap, so the difference is within noise. At xhigh, 25% of answers never arrived: the model used its whole budget thinking.
Work with Leo
Want this for your team?
Leo helps Salesforce teams run AI they own: open-weight models on infrastructure you control, tested on your kind of work before you rely on them. A private benchmark of your shortlist is a fixed fee.
Get a private benchmarkOr email leo@azl.au · azl.au