Studies
Holding everything else still
The leaderboard compares everything with everything. Studies change one thing at a time on the same model, so you can see what that one thing is worth.
012 comparisons
Quantisation
The same model at BF16, FP8 and 4-bit. How much Salesforce ability do you give up to run a model in a quarter of the memory?
Open the study →
021 model
Reasoning effort
Score against effort tier and against output tokens. Where does extra thinking stop paying for itself?
Open the study →
Work with Leo
Want this for your team?
Leo helps Salesforce teams run AI they own: open-weight models on infrastructure you control, tested on your kind of work before you rely on them. A private benchmark of your shortlist is a fixed fee.
Get a private benchmarkOr email leo@azl.au · azl.au