Which models can actually do Salesforce work?

Every row is one config: a model, its quantisation, the engine that served it and the reasoning effort it ran at. Scores are the share of tasks solved on the first try, with the uncertainty drawn in.

Benchmark v0.1.0272 tasks in 15 suites16 configs (4 still running)3 lite effort-sweep runs not ranked hereUpdated 27 September 2026Download JSON

Results table

How to read this. Scores are pass@1 with 95% bootstrap confidence intervals. The dot is the score; the bar is the interval. Overlapping intervals mean no significant difference, so treat configs whose bars overlap as tied. Methodology

Weights

Forcebench leaderboard. Column headers are buttons that sort the table.
1DeepSeek V4.1 FlashEXL3 2.9bpw · vLLM + ExLlamaV3 · highopen weightsEXL3 2.9bpwvLLM + ExLlamaV3high
52%95% CI 47–57
5078718524442874834422671785118%8.2k351 s
2Qwen3.8 Flash-NextNVFP4 · vLLM · mediumopen weightsNVFP4vLLMmedium
34%95% CI 29–39
5044593524622586717173367502%5.3k148 s
3Qwen3.8 Flash-NextMLX 4-bit · MTPLX · mediumopen weightsMLX 4-bitMTPLXmedium
33%95% CI 28–38
25394735011174778224447670615%10k155 s
4Gemma 4 31BQAT W4A16 · vLLM · onopen weightsQAT W4A16vLLMon (max)
30%95% CI 26–35
355059150611747217172766500%2.8k42 s
5Gemma 4 26B-A4BNVFP4 · vLLM · onopen weightsNVFP4vLLMon (max)
28%95% CI 24–33
5039351500176367171127117060%5.9k46 s
6Qwen3.8 27BMLX 8-bit · MTPLX · lowopen weightsMLX 8-bitMTPLXlow
24%95% CI 20–29
353929300017532217224065500%3.2k89 s
7Qwen3.8 27BMLX 4-bit · MTPLX · lowopen weightsMLX 4-bitMTPLXlow
21%95% CI 16–25
303929256617372817172063500%3.5k72 s
8Qwen3.8 27BAWQ-INT4 · vLLM · mediumopen weightsAWQ-INT4vLLMmedium
20%95% CI 17–24
74135154020472015192774802%4.7k35 s
9Qwen3.8 27BSplash 4-bit · Splash · lowopen weightsSplash 4-bitSplashlow
19%95% CI 15–24
103935256011372217112764501%3.5k88 s
10Qwen3.8 27BAWQ-INT4 · vLLM · lowopen weightsAWQ-INT4vLLMlow
19%95% CI 15–23
1028412000647281711400400<1%3.7k27 s
11Qwen3.8 27BAWQ-INT4 · vLLM · xhighopen weightsAWQ-INT4vLLMxhigh (max)
19%95% CI 14–23
033412000114728222213635025%19k149 s
12Qwen3.6 35B-A3BNVFP4 · vLLM · onopen weightsNVFP4vLLMon (max)
18%95% CI 14–22
1533291001111322217171365502%6.1k38 s
–GLM-5.3 FlashEXL3 4.0bpw · vLLM + ExLlamaV3 · maxopen weightspartial2 of 15 suites · 27 of 272 tasks gradedEXL3 4.0bpwvLLM + ExLlamaV3max
45%95% CI 30–60
—73——17——————————0%3.2k180 s
–GLM-5.3 FlashEXL3 4.0bpw · vLLM + ExLlamaV3 · highopen weightspartial14 of 15 suites · 219 of 272 tasks gradedEXL3 4.0bpwvLLM + ExLlamaV3high
40%95% CI 34–46
39394145292225—674541402185150%1.6k92 s
–GLM-5.3 FlashEXL3 4.0bpw · vLLM + ExLlamaV3 · offopen weightspartial2 of 15 suites · 35 of 272 tasks gradedEXL3 4.0bpwvLLM + ExLlamaV3off
37%95% CI 22–51
—56——18——————————0%58236 s
–GLM-5.3 FlashEXL3 4.0bpw · vLLM + ExLlamaV3 · lowopen weightspartial2 of 15 suites · 35 of 272 tasks gradedEXL3 4.0bpwvLLM + ExLlamaV3low
28%95% CI 14–43
—39——18——————————0%23816 s
score and 95% interval suite cells tinted by scoreTokens: mean output tokens per task · Latency: mean seconds per taskpartial not every suite run yet, so not ranked

Want this for your team?

Leo helps Salesforce teams run AI they own: open-weight models on infrastructure you control, tested on your kind of work before you rely on them. A private benchmark of your shortlist is a fixed fee.

Get a private benchmark

Or email leo@azl.au · azl.au