Leaderboard
Which models can actually do Salesforce work?
Every row is one config: a model, its quantisation, the engine that served it and the reasoning effort it ran at. Scores are the share of tasks solved on the first try, with the uncertainty drawn in.
Results table
How to read this. Scores are pass@1 with 95% bootstrap confidence intervals. The dot is the score; the bar is the interval. Overlapping intervals mean no significant difference, so treat configs whose bars overlap as tied. Methodology
| 1 | DeepSeek V4.1 FlashEXL3 2.9bpw · vLLM + ExLlamaV3 · high | EXL3 2.9bpw | vLLM + ExLlamaV3 | high | 52%95% CI 47–57 | 50 | 78 | 71 | 85 | 24 | 44 | 28 | 74 | 83 | 44 | 22 | 67 | 17 | 85 | 11 | 8% | 8.2k | 351 s |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 2 | Qwen3.8 Flash-NextNVFP4 · vLLM · medium | NVFP4 | vLLM | medium | 34%95% CI 29–39 | 50 | 44 | 59 | 35 | 24 | 6 | 22 | 58 | 67 | 17 | 17 | 33 | 6 | 75 | 0 | 2% | 5.3k | 148 s |
| 3 | Qwen3.8 Flash-NextMLX 4-bit · MTPLX · medium | MLX 4-bit | MTPLX | medium | 33%95% CI 28–38 | 25 | 39 | 47 | 35 | 0 | 11 | 17 | 47 | 78 | 22 | 44 | 47 | 6 | 70 | 6 | 15% | 10k | 155 s |
| 4 | Gemma 4 31BQAT W4A16 · vLLM · on | QAT W4A16 | vLLM | on (max) | 30%95% CI 26–35 | 35 | 50 | 59 | 15 | 0 | 6 | 11 | 74 | 72 | 17 | 17 | 27 | 6 | 65 | 0 | 0% | 2.8k | 42 s |
| 5 | Gemma 4 26B-A4BNVFP4 · vLLM · on | NVFP4 | vLLM | on (max) | 28%95% CI 24–33 | 50 | 39 | 35 | 15 | 0 | 0 | 17 | 63 | 67 | 17 | 11 | 27 | 11 | 70 | 6 | 0% | 5.9k | 46 s |
| 6 | Qwen3.8 27BMLX 8-bit · MTPLX · low | MLX 8-bit | MTPLX | low | 24%95% CI 20–29 | 35 | 39 | 29 | 30 | 0 | 0 | 17 | 53 | 22 | 17 | 22 | 40 | 6 | 55 | 0 | 0% | 3.2k | 89 s |
| 7 | Qwen3.8 27BMLX 4-bit · MTPLX · low | MLX 4-bit | MTPLX | low | 21%95% CI 16–25 | 30 | 39 | 29 | 25 | 6 | 6 | 17 | 37 | 28 | 17 | 17 | 20 | 6 | 35 | 0 | 0% | 3.5k | 72 s |
| 8 | Qwen3.8 27BAWQ-INT4 · vLLM · medium | AWQ-INT4 | vLLM | medium | 20%95% CI 17–24 | 7 | 41 | 35 | 15 | 4 | 0 | 20 | 47 | 20 | 15 | 19 | 27 | 7 | 48 | 0 | 2% | 4.7k | 35 s |
| 9 | Qwen3.8 27BSplash 4-bit · Splash · low | Splash 4-bit | Splash | low | 19%95% CI 15–24 | 10 | 39 | 35 | 25 | 6 | 0 | 11 | 37 | 22 | 17 | 11 | 27 | 6 | 45 | 0 | 1% | 3.5k | 88 s |
| 10 | Qwen3.8 27BAWQ-INT4 · vLLM · low | AWQ-INT4 | vLLM | low | 19%95% CI 15–23 | 10 | 28 | 41 | 20 | 0 | 0 | 6 | 47 | 28 | 17 | 11 | 40 | 0 | 40 | 0 | <1% | 3.7k | 27 s |
| 11 | Qwen3.8 27BAWQ-INT4 · vLLM · xhigh | AWQ-INT4 | vLLM | xhigh (max) | 19%95% CI 14–23 | 0 | 33 | 41 | 20 | 0 | 0 | 11 | 47 | 28 | 22 | 22 | 13 | 6 | 35 | 0 | 25% | 19k | 149 s |
| 12 | Qwen3.6 35B-A3BNVFP4 · vLLM · on | NVFP4 | vLLM | on (max) | 18%95% CI 14–22 | 15 | 33 | 29 | 10 | 0 | 11 | 11 | 32 | 22 | 17 | 17 | 13 | 6 | 55 | 0 | 2% | 6.1k | 38 s |
| – | GLM-5.3 FlashEXL3 4.0bpw · vLLM + ExLlamaV3 · max2 of 15 suites · 27 of 272 tasks graded | EXL3 4.0bpw | vLLM + ExLlamaV3 | max | 45%95% CI 30–60 | — | 73 | — | — | 17 | — | — | — | — | — | — | — | — | — | — | 0% | 3.2k | 180 s |
| – | GLM-5.3 FlashEXL3 4.0bpw · vLLM + ExLlamaV3 · high14 of 15 suites · 219 of 272 tasks graded | EXL3 4.0bpw | vLLM + ExLlamaV3 | high | 40%95% CI 34–46 | 39 | 39 | 41 | 45 | 29 | 22 | 25 | — | 67 | 45 | 41 | 40 | 21 | 85 | 15 | 0% | 1.6k | 92 s |
| – | GLM-5.3 FlashEXL3 4.0bpw · vLLM + ExLlamaV3 · off2 of 15 suites · 35 of 272 tasks graded | EXL3 4.0bpw | vLLM + ExLlamaV3 | off | 37%95% CI 22–51 | — | 56 | — | — | 18 | — | — | — | — | — | — | — | — | — | — | 0% | 582 | 36 s |
| – | GLM-5.3 FlashEXL3 4.0bpw · vLLM + ExLlamaV3 · low2 of 15 suites · 35 of 272 tasks graded | EXL3 4.0bpw | vLLM + ExLlamaV3 | low | 28%95% CI 14–43 | — | 39 | — | — | 18 | — | — | — | — | — | — | — | — | — | — | 0% | 238 | 16 s |
Why intervals matter
A few tasks going one way or the other can move a score by several points. The interval shows how much, so a two-point gap between configs isn't sold as a win.
Does quantisation hurt?
The same model at BF16, FP8 and 4-bit, side by side, with the interval on each.
Is more thinking worth it?
Score against reasoning effort, and against the output tokens that effort costs.
Work with Leo
Want this for your team?
Leo helps Salesforce teams run AI they own: open-weight models on infrastructure you control, tested on your kind of work before you rely on them. A private benchmark of your shortlist is a fixed fee.
Get a private benchmarkOr email leo@azl.au · azl.au