Study · Quantisation
Does squeezing a model hurt its Salesforce work?
Quantisation stores a model's weights in fewer bits so it needs less memory and runs cheaper. BF16 is the usual full precision. FP8 halves the size; 4-bit formats like AWQ-INT4 or GGUF Q4_K_M cut it to roughly a quarter. It is how most teams run open models privately. The question is what it costs in quality.
Each chart below is one model at one effort setting, run at different quantisations. Everything else stays fixed: same prompts, same tasks, same sampling settings.
Show the numbers
| Quantisation | Engine | pass@1 | 95% CI | vs MLX 8-bit | No answer | Output tokens | Latency |
|---|---|---|---|---|---|---|---|
| MLX 8-bit | MTPLX | 24% | 20–29 | — | 0% | 3.2k | 89 s |
| MLX 4-bit | MTPLX | 21% | 16–25 | −4 pts | 0% | 3.5k | 72 s |
| Splash 4-bit | Splash | 19% | 15–24 | −5 pts | 1% | 3.5k | 88 s |
| AWQ-INT4 | vLLM | 19% | 15–23 | −5 pts | <1% | 3.7k | 27 s |
Every quantisation lands within noise of MLX 8-bit. The largest gap is −5 pts (AWQ-INT4).
Show the numbers
| Quantisation | Engine | pass@1 | 95% CI | vs MLX 4-bit | No answer | Output tokens | Latency |
|---|---|---|---|---|---|---|---|
| MLX 4-bit | MTPLX | 33% | 28–38 | — | 15% | 10k | 155 s |
| NVFP4 | vLLM | 34% | 29–39 | +1 pt | 2% | 5.3k | 148 s |
Every quantisation lands within noise of MLX 4-bit. The largest gap is +1 pt (NVFP4).
* intervals do not overlap. Everything else is within noise: the gap could plausibly flip on a different sample of tasks. See how the intervals work.
Work with Leo
Want this for your team?
Leo helps Salesforce teams run AI they own: open-weight models on infrastructure you control, tested on your kind of work before you rely on them. A private benchmark of your shortlist is a fixed fee.
Get a private benchmarkOr email leo@azl.au · azl.au