Fifteen corners of the platform

Each suite is a set of tasks about one part of Salesforce development. Most are graded by running the answer; the rest by deterministic validators. None are graded by another model.

272 tasks15 suitesBenchmark v0.1.0
apexExecution

Apex

General Apex engineering: collections and aggregate SOQL, partial-success DML, triggers and handlers (bulkification, recursion), batch, queueable (chaining, finalizers) and schedulable jobs, platform events, callouts through Named Credentials, JSON, dynamic SOQL, Schema describe, and security at API 67.0 (user mode by default). Each answer is deployed check-only to a clean scratch org at API 67.0 and must pass hidden Apex tests, including 200-record bulk tests. Test-writing tasks are graded by mutation testing: the model's tests must pass on the given implementation and fail on every hidden buggy mutant.

20 tasks:6 easy8 medium6 hard
Best: 50% · DeepSeek V4.1 Flash
apiDeterministic validators

Salesforce APIs

Integration work against Salesforce's APIs: REST sObject, query, composite, composite graph and sObject Collections resources, Bulk API 2.0 ingest and query jobs, Tooling API, User Interface API, platform events, custom Apex REST endpoints and OAuth 2.0 token requests. Answers are raw HTTP requests, checked deterministically for method, path, query, headers and body (JSON, form-encoded or CSV), plus a few questions on API behaviour and choice.

18 tasks:6 easy7 medium5 hard
Best: 78% · DeepSeek V4.1 Flash
ciDeterministic validators

CI/CD

Salesforce CI/CD engineering: GitHub Actions workflows and pipeline scripts that authenticate with JWT or SFDX auth URLs from secrets, validate pull requests, quick deploy validated releases, run delta deployments with sfdx-git-delta, manage scratch orgs per pull request, build and promote unlocked packages, gate builds on Code Analyzer and run LWC Jest, plus CI/CD judgement questions. Every sf command in an answer is validated against the pinned sf CLI manifest; workflow properties (triggers, job ordering, environments, cleanup, secrets) are checked deterministically, including simulated events. No org is needed.

17 tasks:5 easy7 medium5 hard
Best: 71% · DeepSeek V4.1 Flash
cliDeterministic validators

Salesforce CLI

Day-to-day and CI/CD work with the Salesforce CLI (sf v2): deploy and retrieve by source, metadata and manifest, validation and quick deploy, destructive changes, Apex tests and anonymous Apex, scratch orgs and sandboxes, JWT auth, config and aliases, data exports and bulk loads, and unlocked packages. Answers are shell commands. Each command is parsed the way the CLI parses it and checked against the command manifest of a pinned @salesforce/cli release (unknown commands or flags, invalid option values, missing required flags and deprecated sfdx-style syntax fail), then compared with the expected commands and flag values.

20 tasks:6 easy8 medium6 hard
Best: 85% · DeepSeek V4.1 Flash
docsDeterministic validators

Salesforce docs

Navigating the official Salesforce documentation: limits, allocations and platform behaviour an engineer has to look up while designing a solution. Each answer is a short, exact fact and must cite the developer.salesforce.com page that states it; both the fact and the cited page are checked deterministically.

17 tasks:5 easy7 medium5 hard
Best: 24% · DeepSeek V4.1 Flash
fflibExecution

fflib Enterprise Patterns

Apex Enterprise Patterns with the open-source fflib libraries (fflib-apex-common and fflib-apex-mocks): Selector, Domain, Service and Unit of Work layers, the Application factory class, and unit tests with ApexMocks. Code is deployed to a scratch org with the pinned fflib commits installed and must pass hidden Apex tests; test-writing tasks are graded by mutation testing (the model's tests must pass against the correct code and fail against seeded bugs).

18 tasks:5 easy7 medium6 hard
Best: 44% · DeepSeek V4.1 Flash
flowExecution

Flow

Salesforce Flow engineering: record-triggered flows (before-save fast field updates, after-save related-record automation, before-delete), entry conditions and $Record__Prior, decisions and formulas, Custom Error, fault paths, subflows, autolaunched flows called from Apex, invocable Apex actions, scheduled paths, and choosing between flows and Apex triggers. Answers are Flow metadata files (API 67.0) deployed check-only to a clean scratch org, where hidden Apex tests do DML (including 200-record bulk tests that assert on Limits) and check the outcome.

18 tasks:5 easy7 medium6 hard
Best: 28% · DeepSeek V4.1 Flash
limitsExecution

Governor limits & pushback

Governor limits in patterns, and the judgement to push back. Most tasks are briefs in which a stakeholder explicitly asks for something that breaks at scale (a query, DML statement, callout, email, platform event publish or async job per record, an unfiltered query, a quadratic loop). The answer is deployed check-only to a clean scratch org at API 67.0 and must pass hidden Apex tests that use 200 records or more and assert both the functional outcome and `Limits` usage, so following the brief literally fails; the reply must also tell the user why it deviated (checked with per-task patterns on the prose and code comments). Control tasks make requests that look similar but are fine and must simply be implemented. Analysis tasks ask how many queries or DML statements code consumes (helper methods, trigger chunks of 200, cascading triggers), which limit breaks first, and how synchronous and asynchronous limits differ.

19 tasks:6 easy7 medium6 hard
Best: 74% · DeepSeek V4.1 Flash
lwcExecution

Lightning Web Components

Lightning Web Components engineering: public properties and reactivity, templates and conditional rendering, events, parent/child APIs, wire adapters, imperative Apex, toasts, navigation, Lightning Message Service, lightning-datatable, forms, lifecycle cleanup, slots, labels, App Builder metadata and accessibility. Each answer runs under hidden Jest tests (@salesforce/sfdx-lwc-jest); no org is needed.

18 tasks:5 easy8 medium5 hard
Best: 83% · DeepSeek V4.1 Flash
npspExecution + validators

Nonprofit Success Pack

Engineering on the Nonprofit Success Pack (NPSP) managed packages (npsp, npe01, npo02, npe03, npe4, npe5): the Household Account model, gifts and payments, GAU allocations, soft credits, Enhanced Recurring Donations, affiliations and relationships, TDTM trigger handlers and the BDI data import API. Code answers are deployed to a scratch org with NPSP installed and must pass hidden Apex tests that exercise NPSP's own automation; SOQL answers run against seeded NPSP data; knowledge questions cover NPSP architecture and configuration.

18 tasks:5 easy8 medium5 hard
Best: 44% · DeepSeek V4.1 Flash
packagingDeterministic validators

Packaging

Package development with unlocked, second-generation (2GP) and first-generation (1GP) managed packages: sfdx-project.json package directories, aliases, dependencies, ancestry and test setup; `sf package` and `sf package1` commands for creating, versioning, promoting, installing, upgrading and push-upgrading packages; and the platform rules behind them (IDs, promotion requirements, upgrade paths, component removal). Graded deterministically: project files are resolved the way the CLI resolves them, commands are validated against the pinned `sf` command manifest, knowledge questions are multiple choice.

18 tasks:5 easy7 medium6 hard
Best: 44% · Qwen3.8 Flash-Next
permissionsExecution + validators

Permissions & access

Least-privilege access design: permission sets (object, field, tab, app, Apex class, user and custom permissions), session-based permission sets, permission set groups with muting, fixing over-permissive access, and Apex that checks access (custom permissions, stripInaccessible, UserRecordAccess) under API 67.0's user-mode default. Metadata answers are deployed check-only to a scratch org with hidden custom objects; hidden Apex tests create Minimum Access users, assign the answer's permission sets or groups and assert effective access with System.runAs. A few knowledge tasks cover how profiles, permission sets, groups, muting and sharing combine.

15 tasks:4 easy6 medium5 hard
Best: 67% · DeepSeek V4.1 Flash
scratch-defExecution + validators

Scratch org definitions

Turn a client brief into a working scratch org definition file (and sfdx-project.json). Answers are validated against the official JSON schemas, the documented scratch org feature list and the Metadata API settings types; settings are check-only deployed to a dedicated grader org exactly as scratch org creation applies them (types with lasting side effects, such as fiscal year or currencies, are validated offline). Task rules then check the brief's requirements (edition, features, settings values, locale).

18 tasks:5 easy7 medium6 hard
Best: 17% · DeepSeek V4.1 Flash
soqlExecution

SOQL

Text-to-SOQL on a realistic B2B sales and service org: filters, relationship queries in both directions, semi- and anti-joins, aggregates, date and fiscal functions, polymorphic relationships, multi-select picklists and null handling. Each answer is executed in a seeded scratch org (orgs/base) and its result set must equal the gold query's (execution accuracy).

20 tasks:6 easy8 medium6 hard
Best: 85% · DeepSeek V4.1 Flash
tafExecution

Trigger Actions Framework

Real work in orgs that use the open-source Trigger Actions Framework (TAF) by Mitch Spano: writing trigger action classes, registering them with custom metadata, ordering, bypass mechanisms, entry criteria, recursion control, DML finalizers, flow actions, migrating legacy triggers and testing actions. Each answer is deployed (check-only) with hidden Apex tests to a scratch org that has TAF 0.3.4 installed; the tests perform real DML and assert that the framework ran the model's actions as specified.

18 tasks:6 easy7 medium5 hard
Best: 11% · DeepSeek V4.1 Flash

Want this for your team?

Leo helps Salesforce teams run AI they own: open-weight models on infrastructure you control, tested on your kind of work before you rely on them. A private benchmark of your shortlist is a fixed fee.

Get a private benchmark

Or email leo@azl.au · azl.au