Sovereign AI
AI you own
I focus on AI models my clients own: open weights, running on infrastructure they control. Forcebench measures exactly those models on Salesforce work, and I can run it for your team for a fixed fee.
Why own the model
Sovereign AI means the model is yours
Open-weight models run on hardware you control, on-premises or in your own cloud account. Nothing about how they behave depends on a vendor’s roadmap.
Your data stays yours
Prompts, code and customer records never leave infrastructure you control. No third-party retention policy to vet, and no data-residency question to answer for every new tool.
It can’t change under you
Hosted models get updated, repriced and retired on someone else’s schedule. Weights you own behave next year exactly as they did when you tested them.
Costs you can plan
You pay for hardware or a fixed hosting bill, not per token. A busy month doesn’t turn into a surprise invoice.
You can check it
You can test the exact weights, quantisation and settings you’ll run in production, and test them again after every change. That is what Forcebench is for.
The catch
Which model, at what size?
Open-weight models come in dozens of sizes, each in several quantisations, served by different engines at different reasoning efforts. A smaller quantisation runs on cheaper hardware, but what does it cost you on real work? Generic benchmarks can’t say, because they don’t test Salesforce.
That is why every model on the Forcebench leaderboard so far is open weight and runs locally, with its quantisation, engine and effort recorded. Right now DeepSeek V4.1 Flash leads at 52%, running at EXL3 2.9bpw. The quantisation study and the reasoning effort study show what smaller builds and extra thinking actually buy you.
Private benchmark
Your shortlist, your kind of work, a fixed fee
Before you buy hardware or commit to a model, find out how your shortlist does on the Salesforce work your team actually does. I run the benchmark and report back.
- 1
Shortlist
We agree a shortlist of models you could realistically own and run, typically three to five, at the quantisations and reasoning efforts that fit your hardware and budget. A hosted model can join as a baseline.
- 2
Test on your kind of work
I run the Forcebench suites that match what your team does, and can add private tasks written from examples of your real work. Same harness as the public leaderboard: check-only deploys to scratch orgs I create, hidden tests, error bars.
- 3
Report back
A written report you can share internally: scores with confidence intervals per suite, speed and token cost, examples of where each model fails, and a clear recommendation of which model, quantisation and effort to run.
Fixed fee, agreed up front. Your tasks and results stay private.
Questions
Good to know
Do you need access to our Salesforce org?
No. Everything runs in scratch orgs created for the benchmark, inside a sandbox that refuses to touch any other org. For private tasks you share examples of the work, anonymised if needed, and I turn them into tasks with hidden tests.
Which models can you test?
Any open-weight model served behind an OpenAI-compatible API (vLLM, SGLang, llama.cpp, MLX and similar), at the quantisations you would actually run. Hosted models from the big labs can be included as a baseline, so you can see what owning your model costs you, if anything.
Will our results be published?
No. Private benchmarks and private tasks stay private unless you ask for them to be published.
Can we run Forcebench ourselves?
Yes. It is open source and free to use (Apache-2.0 code, CC BY 4.0 tasks); you just pay for your own GPUs, plus a Salesforce Dev Hub for the scratch orgs (a free Developer Edition org works). The fixed-fee benchmark is for teams that want it done properly without the setup: choosing candidates, serving each one with the right settings, writing private tasks and interpreting the results.
What does it cost?
A fixed fee, quoted up front. It depends on how many models you want tested and the kinds of tasks. Get in touch for a quote.
Anything else? Email leo@azl.au.