Apex

General Apex engineering: collections and aggregate SOQL, partial-success DML, triggers and handlers (bulkification, recursion), batch, queueable (chaining, finalizers) and schedulable jobs, platform events, callouts through Named Credentials, JSON, dynamic SOQL, Schema describe, and security at API 67.0 (user mode by default). Each answer is deployed check-only to a clean scratch org at API 67.0 and must pass hidden Apex tests, including 200-record bulk tests. Test-writing tasks are graded by mutation testing: the model's tests must pass on the given implementation and fail on every hidden buggy mutant.

20
tasks
Execution
grading
50%
best pass@1 · DeepSeek V4.1 Flash and 2 tied
13
configs with results (1 partial)

How it's graded

Each answer is deployed as a check-only validation to a clean scratch org together with hidden Apex tests, including 200-record bulk tests and governor-limit assertions. It passes only if it compiles and every test passes. Test-writing tasks are graded by mutation testing: the model’s tests must pass against correct code and fail against seeded bugs.

Difficulty mix

6 easy8 medium6 hard

Who's best at Apex

pass@1 on this suite only, with 95% confidence intervals. With 20 tasks the intervals are wide, so read overlapping bars as ties. Open in the full leaderboard →

Apex: pass@1 by config with 95% confidence intervals
#Model and configScore pass@1 · 95% CITasks
1DeepSeek V4.1 Flash EXL3 2.9bpw · vLLM + ExLlamaV3 · effort high
50%95% CI 30–70
20
1Gemma 4 26B-A4B NVFP4 · vLLM · effort on
50%95% CI 30–70
20
1Qwen3.8 Flash-Next NVFP4 · vLLM · effort medium
50%95% CI 30–70
20
4Gemma 4 31B QAT W4A16 · vLLM · effort on
35%95% CI 15–55
20
4Qwen3.8 27B MLX 8-bit · MTPLX · effort low
35%95% CI 15–55
20
6Qwen3.8 27B MLX 4-bit · MTPLX · effort low
30%95% CI 10–50
20
7Qwen3.8 Flash-Next MLX 4-bit · MTPLX · effort medium
25%95% CI 5–45
20
8Qwen3.6 35B-A3B NVFP4 · vLLM · effort on
15%95% CI 0–30
20
9Qwen3.8 27B Splash 4-bit · Splash · effort low
10%95% CI 0–25
20
9Qwen3.8 27B AWQ-INT4 · vLLM · effort low
10%95% CI 0–25
20
11Qwen3.8 27B AWQ-INT4 · vLLM · effort medium
7%95% CI 2–13
20
12Qwen3.8 27B AWQ-INT4 · vLLM · effort xhigh
0%95% CI 0–0
20
–GLM-5.3 Flash partialEXL3 4.0bpw · vLLM + ExLlamaV3 · effort highRun so far: 14 of 15 suites · 219 of 272 tasks graded
39%95% CI 17–61
18

What's in the suite

Solve rate is the mean pass@1 on that task across every finished run, a rough guide to how hard models find it.

Apex tasks with difficulty and solve rate
TaskDifficultySolve rate
Active picklist values via Schema describeapex-picklist-describeeasy
17%
Bulk-safe invocable action for Flowapex-invocable-contact-lookupeasy
25%
Closed-won amount per account with aggregate SOQLapex-won-amount-by-accounteasy
78%
Group contacts by email domainapex-contact-email-domaineasy
53%
Nightly schedulable log purge with retention rulesapex-schedulable-log-purgeeasy
25%
Payment terms parser with a custom exception hierarchyapex-custom-exception-hierarchyeasy
0%
Bulkify a legacy Contact trigger without changing behaviourapex-bulkify-legacy-triggermedium
53%
Parse an order payload whose keys are Apex reserved wordsapex-json-reserved-keysmedium
17%
Partial-success lead import with per-row errorsapex-partial-success-lead-importmedium
42%
Publish and subscribe to shipment platform eventsapex-platform-event-shipmentsmedium
25%
Queueable with a transaction finalizer that logs and retriesapex-queueable-finalizer-retrymedium
8%
REST callout through a Named Credential with a single retryapex-callout-named-credential-retrymedium
36%
Stateful batch that flags overdue invoices and reports totalsapex-batch-stateful-overduemedium
25%
Write tests that catch bugs in a shipping calculatorapex-mutation-tests-shippingmedium
58%
Chained Queueable that is testable end to endapex-queueable-chain-depthhard
0%
Injection-safe dynamic SOQL search in user modeapex-dynamic-soql-user-modehard
17%
Open-pipeline rollup trigger that survives recursionapex-rollup-trigger-recursionhard
17%
Upgrade a service to API 67.0 without changing its security behaviourapex-api67-access-modeshard
0%
Write tests that catch bugs in an Opportunity trigger handlerapex-mutation-tests-won-handlerhard
0%
Write tests with an HttpCalloutMock that pin a REST contractapex-mutation-tests-callouthard
33%

Want this for your team?

Leo helps Salesforce teams run AI they own: open-weight models on infrastructure you control, tested on your kind of work before you rely on them. A private benchmark of your shortlist is a fixed fee.

Get a private benchmark

Or email leo@azl.au · azl.au