Holding everything else still

The leaderboard compares everything with everything. Studies change one thing at a time on the same model, so you can see what that one thing is worth.

Want this for your team?

Leo helps Salesforce teams run AI they own: open-weight models on infrastructure you control, tested on your kind of work before you rely on them. A private benchmark of your shortlist is a fixed fee.

Get a private benchmark

Or email leo@azl.au · azl.au