Does squeezing a model hurt its Salesforce work?

Quantisation stores a model's weights in fewer bits so it needs less memory and runs cheaper. BF16 is the usual full precision. FP8 halves the size; 4-bit formats like AWQ-INT4 or GGUF Q4_K_M cut it to roughly a quarter. It is how most teams run open models privately. The question is what it costs in quality.

Each chart below is one model at one effort setting, run at different quantisations. Everything else stays fixed: same prompts, same tasks, same sampling settings.

Qwen3.8 27B · effort low
Overall pass@1 with 95% confidence interval. Differences are against the highest-precision run.
Qwen3.8 27B at effort low: overall pass@1 by quantisationMLX 8-bit · MTPLX (reference): 24% · 20–29; MLX 4-bit · MTPLX: 21% · 16–25 · −4 pts; Splash 4-bit · Splash: 19% · 15–24 · −5 pts; AWQ-INT4 · vLLM: 19% · 15–23 · −5 pts0%25%50%75%100%MLX 8-bit · MTPLX (reference)24% · 20–29MLX 4-bit · MTPLX21% · 16–25 · −4 ptsSplash 4-bit · Splash19% · 15–24 · −5 ptsAWQ-INT4 · vLLM19% · 15–23 · −5 pts
Show the numbers
QuantisationEnginepass@195% CIvs MLX 8-bitNo answerOutput tokensLatency
MLX 8-bitMTPLX24%20–29—0%3.2k89 s
MLX 4-bitMTPLX21%16–25−4 pts0%3.5k72 s
Splash 4-bitSplash19%15–24−5 pts1%3.5k88 s
AWQ-INT4vLLM19%15–23−5 pts<1%3.7k27 s

Every quantisation lands within noise of MLX 8-bit. The largest gap is −5 pts (AWQ-INT4).

Qwen3.8 Flash-Next · effort medium
Overall pass@1 with 95% confidence interval. Differences are against the highest-precision run.
Qwen3.8 Flash-Next at effort medium: overall pass@1 by quantisationMLX 4-bit · MTPLX (reference): 33% · 28–38; NVFP4 · vLLM: 34% · 29–39 · +1 pt0%25%50%75%100%MLX 4-bit · MTPLX (reference)33% · 28–38NVFP4 · vLLM34% · 29–39 · +1 pt
Show the numbers
QuantisationEnginepass@195% CIvs MLX 4-bitNo answerOutput tokensLatency
MLX 4-bitMTPLX33%28–38—15%10k155 s
NVFP4vLLM34%29–39+1 pt2%5.3k148 s

Every quantisation lands within noise of MLX 4-bit. The largest gap is +1 pt (NVFP4).

* intervals do not overlap. Everything else is within noise: the gap could plausibly flip on a different sample of tasks. See how the intervals work.

Want this for your team?

Leo helps Salesforce teams run AI they own: open-weight models on infrastructure you control, tested on your kind of work before you rely on them. A private benchmark of your shortlist is a fixed fee.

Get a private benchmark

Or email leo@azl.au · azl.au