Write a task
A good Forcebench task is small, self-contained, and gradable by running it. Each one has:
- A prompt: what to build or answer, plus any starting metadata or code, written the way you'd brief a colleague.
- A grader: hidden Apex tests, Jest tests, or a deterministic validator. The model never sees it.
- A reference answer, written exactly as a model would reply, that passes the grader.
- Alternative answers: other correct solutions in a different style, which must also pass.
- Negative answers: plausible, common mistakes, which must fail.
- Metadata: suite, difficulty (easy, medium or hard) and the canary GUID.
Before a task is accepted it has to pass oracle validation(forcebench validate, which CI runs on every change): the reference and alternative answers pass, and the negative answers and an empty reply fail. The prompt has to state everything the grader checks, so any correct solution passes. If you can't think of a wrong answer the grader would catch, the grader probably needs more tests.
Some things that make a task better:
- It reflects something that bites in real orgs: governor limits, bulk data, sharing, order of execution, package quirks.
- There's one clear definition of done, even if there are many ways to get there.
- It's written from scratch. Please don't copy tasks, answers or text from Trailhead, Stack Exchange or paid courses.
The repo's contributing guide and task-authoring guide have the current file layout and the commands to validate a task locally. A private held-out split is planned for v0.2; those tasks will be handled separately so they stay private. If you'd like to write some, email leo@azl.au.
Run a model
The harness runs any config you can serve: a hosted provider, or an open-weights model behind any OpenAI-compatible endpoint (vLLM, SGLang, LM Studio and so on). A config is defined by four things, and all four are reported:
- Model, as the vendor names it
- Quantisation, e.g. BF16, FP8, AWQ-INT4, GGUF Q4_K_M
- Inference engine, e.g. vLLM, SGLang, llama.cpp, or the vendor's API
- Reasoning effort, in the vendor's own terms
Use the vendor's recommended sampling settings and the standard prompts; results with custom prompts can't go on the leaderboard. Open a pull request with the full run directory under results/runs/. Submitted results are marked self-reported until a maintainer reproduces them.
Report a problem
Every answer and grade is public, so if a result looks wrong you can check what the model actually wrote. If a grader got it wrong (a correct answer fails, or a wrong one passes), open an issue with the task id and the answer. Fairness bugs get priority. A fixed task gets a new version, and stored answers can be re-graded without calling the model again.
Licence
Forcebench code is Apache-2.0 and the data (tasks, results and published answers) is CC BY 4.0. Contributions are accepted under the same licences. Use the results however you like; just credit Forcebench.