The short version
- Models get real Salesforce tasks across fifteen suites. Their answers are run wherever possible: deployed as a check-only validation to a clean scratch org with hidden tests, run under Jest, executed against seeded data, or checked by deterministic validators. No AI model grades another.
- The score is pass@1: the share of tasks a config gets right on its one attempt. v0.1 runs one sample per task; repeated samples and pass@k come in v0.2.
- Every score comes with a 95% confidence interval. Overlapping intervals mean no significant difference.
- Every model gets the same prompt, word for word (one turn, no tools), and runs with its vendor's recommended sampling settings and a 32,768-token output budget.
- Reasoning effort and quantisation are recorded as part of each config, not hidden.
- Answers are generated first, straight from the inference server with no hidden retries, then graded in a sandbox. Answers that time out are re-run, not scored.
- Headline results use the full task set. Effort sweeps may use a fixed lite subset, which is only ever compared with other lite results.
- Every task's reference answer must pass its grader and known-wrong answers must fail it. Tasks carry a canary GUID, and every answer and grade is published. A private held-out split is planned for v0.2.
Run it, don’t judge it
The benchmarks that have held up best in AI coding, like HumanEval and SWE-bench, share one idea: a task passes when the code works, not when it looks right. Forcebench follows the same rule wherever it can.
- Check-only deploys with hidden tests. Apex, Trigger Actions Framework, fflib, Flow, NPSP code and permissions answers are deployed as a check-only validation to a clean scratch org, together with hidden Apex tests, including 200-record bulk tests and governor-limit assertions. The tests have to pass, the way a reviewer's test class would exercise your code. If the answer doesn't compile or a test fails, it scores zero. Where the task is to write tests, mutation testing checks that the model's tests pass against correct code and fail against seeded bugs.
- Effective access. For permission sets and groups, hidden tests create Minimum Access users, assign the answer's permissions and assert what those users can actually read, edit and do, using
System.runAs. - Jest. Lightning Web Components run under hidden Jest tests (
@salesforce/sfdx-lwc-jest). - Execution accuracy. SOQL answers run against a scratch org seeded with fixed data, and the result set must equal the gold query's, as in the BIRD and Spider text-to-SQL benchmarks.
- Deterministic validators. Some answers can't simply be run: a CLI command, a packaging config, a GitHub Actions pipeline, an HTTP request. These are parsed and checked by code. Every
sfcommand is checked against the command manifest of a pinned Salesforce CLI release, so an invented or deprecated flag fails. Scratch org definitions are also validated against the official schemas and feature list, and their settings are deployed check-only. - Docs. Answers are short, exact facts, and the cited documentation page must be right too.
No AI model grades any v0.1 task. Model-as-judge grading is cheap and flexible, but it's biased towards confident, well-formatted answers, and it can't tell you whether a trigger survives a bulk insert. Where a task is about knowledge or judgement (before-save flow, after-save flow or Apex?), it's multiple choice or has a short exact answer.
Answer extraction is forgiving about prose and strict about the answer: explain as much as you like, but if the command block, file or Answer: line is missing, the task fails.
Suite by suite
| Suite | Grading | How |
|---|---|---|
| Apex | Execution | Each answer is deployed as a check-only validation to a clean scratch org together with hidden Apex tests, including 200-record bulk tests and governor-limit assertions. It passes only if it compiles and every test passes. Test-writing tasks are graded by mutation testing: the model’s tests must pass against correct code and fail against seeded bugs. |
| Salesforce APIs | Deterministic validators | Answers are raw HTTP requests, checked deterministically for method, path, query, headers and body (JSON, form-encoded or CSV). A few questions cover API behaviour and choosing the right API. |
| CI/CD | Deterministic validators | GitHub Actions workflows are checked for structure and secrets hygiene, and every sf command in them is validated against the pinned sf command manifest. |
| Salesforce CLI | Deterministic validators | Each command is parsed the way the CLI parses it and checked against the command manifest of a pinned Salesforce CLI release: unknown commands or flags, invalid values, missing required flags and deprecated sfdx-style syntax fail. It is then compared with the expected commands and flag values. |
| Salesforce docs | Deterministic validators | Each answer is a short, exact fact and must cite the developer.salesforce.com page that states it. Both the fact and the cited page are checked deterministically. |
| fflib Enterprise Patterns | Execution | Code is deployed as a check-only validation to a scratch org with the pinned fflib libraries installed and must pass hidden Apex tests. ApexMocks test-writing tasks are graded by mutation testing: the model’s tests must pass against the correct code and fail against seeded bugs. |
| Flow | Execution | Flow metadata is deployed as a check-only validation to a clean scratch org, where hidden Apex tests do DML (including 200-record bulk tests that assert on governor limits) and check the outcome. Decision-guide questions (flow or Apex?) are multiple choice. |
| Governor limits & pushback | Execution | Each task asks for something that invites an anti-pattern, such as SOQL inside a loop. The code is deployed as a check-only validation with hidden bulk tests that assert on governor limits, and the reply must also flag the problem instead of quietly doing what was asked. |
| Lightning Web Components | Execution | Each component runs under hidden Jest tests (@salesforce/sfdx-lwc-jest) that render it and check its behaviour. No org is needed. |
| Nonprofit Success Pack | Execution + validators | Code answers are deployed as a check-only validation to a scratch org with NPSP installed and must pass hidden Apex tests that exercise NPSP’s own automation. SOQL answers run against seeded NPSP data; knowledge questions cover NPSP architecture and configuration. |
| Packaging | Deterministic validators | sfdx-project.json files are resolved the way the CLI resolves them, sf package and sf package1 commands are validated against the pinned sf command manifest, and packaging-judgement questions (1GP, managed 2GP or unlocked) are multiple choice. |
| Permissions & access | Execution + validators | Permission sets and groups are deployed as a check-only validation to a scratch org. Hidden Apex tests create Minimum Access users, assign the answer’s permission sets or groups and assert effective access with System.runAs. A few knowledge tasks are multiple choice or have a short exact answer. |
| Scratch org definitions | Execution + validators | Definition files are validated against the official JSON schemas and the documented scratch org feature list, and their settings are deployed check-only to a grader org exactly as scratch org creation applies them. Task rules then check the brief: edition, features, settings values and locale. |
| SOQL | Execution | Each query runs against a scratch org seeded with fixed data, and its result set must equal the gold query’s (execution accuracy, as in BIRD and Spider). A query that errors or returns different rows fails. |
| Trigger Actions Framework | Execution | Each answer is deployed as a check-only validation, with hidden Apex tests, to a scratch org with the Trigger Actions Framework installed. The tests do real DML and assert that the framework ran the model’s actions as specified: ordering, bypasses, entry criteria and recursion control. |
What a score means: pass@1
In v0.1 each config answers each task once. A task scores 1 if that answer passes its grader and 0 if it doesn't. A suite score is the mean over the suite's tasks, and the overall score is the average of the suite scores, so a suite with more tasks doesn't dominate.
A score of 62% means: pick a suite and a task in it, ask the model once and use what comes back, and it works about 62% of the time. That's the number that matters when a developer asks a model once and uses the answer.
Models aren't deterministic: ask the same one twice and you can get one answer that works and one that doesn't. With one sample per task, any single task's result is partly luck, which is why Forcebench only reports scores across many tasks, with intervals. From v0.2, runs will take repeated samples (n ≥ 3 per task) and count how many pass (c). pass@1 then becomes the mean of c / n, and the same samples givepass@k, the chance that at least one of k attempts works, via the standard unbiased estimator from the HumanEval paper:
pass@k = mean over tasks of ( 1 − C(n − c, k) / C(n, k) )where C(a, b) is "a choose b". For k = 1 this reduces to c / n.
Error bars: 95% bootstrap confidence intervals
With tens of tasks per suite rather than thousands, a score is an estimate. Swap in a slightly different set of equally fair tasks and it would move. The confidence interval says by how much.
Forcebench uses a bootstrap: take the task list, draw a new list of the same length by picking tasks at random with replacement (so some appear twice and some not at all), recompute the score, and do that 10,000 times. The middle 95% of those recomputed scores is the interval. For the overall score, tasks are resampled within each suite, so every resample keeps the same suite mix.
The resampling is over tasks, not individual answers. With one answer per task in v0.1 that's the same thing. It matters from v0.2: several attempts at the same task aren't independent pieces of evidence (a model with a blind spot for polymorphic SOQL will probably miss every attempt), and treating them as independent would make the intervals look narrower, and the results more certain, than they really are.
How to read them: if two configs' intervals overlap, Forcebench doesn't claim either is better. That rule is deliberately conservative. Two configs whose intervals overlap slightly can still differ in a careful paired test on the same tasks, but the leaderboard won't call a winner on that basis. When the intervals don't overlap, the difference is real for this benchmark.
Per-suite intervals are much wider than overall ones because each suite has fewer tasks. Treat a per-suite ranking as a hint, the overall ranking as the evidence.
One prompt for every model
Every model gets the same system prompt and the same fixed output-format instructions for each kind of answer, word for word. There are no worked examples (it's zero-shot), no tools and no per-model prompt tuning: one user turn, one reply.
Some models would score higher with a prompt written just for them. That's true and it's the point: Forcebench measures how a model does with a reasonable, fixed instruction, which is closer to how most developers use them. The only thing that differs is the chat template, which the inference engine applies the way each model expects.
Sampling settings
Each model runs with its vendor's recommended sampling settings for thinking mode (temperature, top-p and so on), and the exact request fields are published with every run. Forcing every model to the same settings sounds fairer but isn't: temperature 0 is not a neutral default, and some reasoning models do noticeably worse, or get stuck repeating themselves, when run greedily. Using the settings the vendor tested with measures each model as it's meant to be used.
Reasoning effort is a dimension, not a footnote
Many models can be told how hard to think before answering. The same model at low and high effort can differ by more than two different models do, and uses very different amounts of time and tokens. So Forcebench treats the same model at a different effort as a different config with its own row.
Vendors name effort differently (low, medium, high, xhigh, thinking on or off). Each result records the vendor's own setting and a normalised tier (off, low, medium, high, max) for comparison, along with mean output tokens, reasoning tokens and latency per task. The effort study shows what extra thinking buys.
Output budget, not a clock
Every answer gets an output budget of 32,768 tokens, reasoning included. A configuration that thinks past it without producing an answer fails the task. That's a real cost of running a model at that effort, so it counts against the score. It's the same length DeepSeek-R1 was evaluated with and Qwen3 recommends for most tasks.
The leaderboard reports these separately as the No answer rate, so "answered wrongly" and "never answered" can be told apart. The reasoning a model wrote before running out is kept with the run's published replies, so anyone can check whether it was working through the problem or going in circles.
The budget is tokens, never wall-clock time: a slow machine mustn't cost a model points. A request that times out is treated like an endpoint failure and re-run, not scored.
How a run executes
Runs follow the shape of established code benchmarks such as SWE-bench, EvalPlus and LiveCodeBench, in two phases:
- Generate. Every answer is requested from the model and stored; nothing else happens in this phase. Requests go directly to the inference server (vLLM, SGLang and so on) or the vendor's API, never through a general-purpose proxy whose timeouts, retries or defaults would become part of the measurement. Responses are streamed, and the exact request fields (sampling, effort switch, token budget) are recorded. Interrupted runs resume without redoing finished answers.
- Grade. Stored answers are graded in a sandbox: check-only deploys with hidden tests, Jest, SOQL execution and deterministic validators. Grading can be repeated at any time without calling the model, for example after a grader fix.
No hidden retries. Client-side retries are off. A proxy or SDK that silently restarts slow requests would keep only the answers that happened to finish quickly, biasing slow configurations towards short answers. Answers affected before this was fixed have been invalidated and regenerated, with the history kept.
Full and lite. Headline results use the full task set. Expensive sweeps, like one model at several effort levels, may use lite: a fixed, stratified subset of 4 tasks per suite (1 easy, 2 medium, 1 hard), chosen by a hash of the task id. Lite results are only compared with other lite results, are labelled wherever they appear, and never enter the leaderboard.
Quantisation and inference engines are recorded
An open-weights model isn't one thing. The same model can be served at full precision (BF16), at 8 bits (FP8) or at 4 bits (AWQ-INT4, GGUF Q4_K_M and others), by different inference engines such as vLLM, SGLang or llama.cpp. Each can change the results, so every config records its model, quantisation, engine and effort.
Closed models are run through their vendor's API. Their precision isn't public, so they appear as "vendor-served". The quantisation study compares the same model across precisions.
Every task is tested too
A grader that passes everything, or nothing, is worse than no grader. So every task ships with:
- a reference answer, written exactly as a model would reply, that must pass its grader;
- alternative answers, correct solutions in a different style, that must also pass, so a grader can't quietly accept only one phrasing;
- negative answers, plausible common mistakes (SOQL in a loop, the wrong sharing keyword, a deprecated
sfdx force:command, an invalid scratch org feature), that must fail; - and an empty reply, which must fail.
forcebench validate runs all of these, and CI runs it on every change. Prompts state everything the grader checks (names, signatures, paths, return shapes), so any correct solution passes. It's the same reason you'd want a test class that fails when the trigger is commented out.
Isolation: nothing is committed
Answers that need an org are deployed as check-only validations, together with the hidden tests, to long-lived grader scratch orgs. A check-only deploy compiles the code and runs the tests, then rolls everything back, so nothing is ever committed. Every attempt starts from the same org state, and attempts can't affect each other or later runs. SOQL tasks query a scratch org seeded with fixed data.
Salesforce enforces 75% per-class code coverage on validations that run specified tests, even in scratch orgs. Forcebench ignores those coverage warnings: only compile results and test outcomes count.
Contamination controls
If a model has seen the tasks during training, its score measures memory, not skill. Forcebench guards against that in a few ways:
- A canary GUID in every task. Each task file carries the Forcebench canary string (the BIG-bench convention). Model trainers who respect canaries can filter the files out of training data, and a model that can reproduce the string has almost certainly seen them.
- Original tasks. Tasks are written for Forcebench, never copied from certification exams, Trailhead or blogs.
- Dated tasks. Each task records when it was created, so results can be sliced by task age.
- A private held-out split, from v0.2. A set of tasks written the same way but never published will run alongside the public tasks (the harness can already load extra private suites). A model that does much better on the public tasks than on the held-out ones will be flagged.
Published answers
Every run publishes each final answer, its grade and its token counts in the benchmark repo (results/runs/<run>/cases.jsonl), and the full replies, including reasoning where the provider returns it, as a release download. Each run also records the benchmark version, the harness commit, the task versions, the config and the exact request fields. If a number looks wrong, you can read exactly what the model wrote and why it passed or failed. Stored answers can be re-graded without calling the model again, so a grader fix applies to past runs. Found a grading bug? Open an issue.
Versioning
Forcebench is versioned; this site is built from v0.1.0, with fifteen suites. Scores are only comparable within the same benchmark version. Changing a task bumps that task's version, and results from older versions of a task are excluded from the leaderboard.
What Forcebench doesn’t tell you
- It's single-turn with no tools: it measures what a model knows and can produce in one go, not how well it works as an agent inside a repo with an org. An agentic track is on the roadmap.
- Tasks are small and self-contained so they can be graded automatically. Real projects are messier, with more context, legacy code and people.
- It can't tell you how a model will do on your org, your naming conventions or your managed packages.
- Scratch orgs are Developer edition unless a suite needs otherwise, so some enterprise features are validated structurally rather than in an org of that edition.
- Salesforce changes three times a year. Tasks pin the API version they rely on, and the
sfcommand manifest is pinned and refreshed deliberately. - Suites have tens of tasks, not thousands, and v0.1 takes one sample per task. The intervals say so honestly; read them.
If you need those answers for your own team's work, that's what Leo does at Azul Labs.