Governor limits & pushback

Governor limits in patterns, and the judgement to push back. Most tasks are briefs in which a stakeholder explicitly asks for something that breaks at scale (a query, DML statement, callout, email, platform event publish or async job per record, an unfiltered query, a quadratic loop). The answer is deployed check-only to a clean scratch org at API 67.0 and must pass hidden Apex tests that use 200 records or more and assert both the functional outcome and `Limits` usage, so following the brief literally fails; the reply must also tell the user why it deviated (checked with per-task patterns on the prose and code comments). Control tasks make requests that look similar but are fine and must simply be implemented. Analysis tasks ask how many queries or DML statements code consumes (helper methods, trigger chunks of 200, cascading triggers), which limit breaks first, and how synchronous and asynchronous limits differ.

19
tasks
Execution
grading
74%
best pass@1 · DeepSeek V4.1 Flash and 1 tied
12
configs with results

How it's graded

Each task asks for something that invites an anti-pattern, such as SOQL inside a loop. The code is deployed as a check-only validation with hidden bulk tests that assert on governor limits, and the reply must also flag the problem instead of quietly doing what was asked.

Difficulty mix

6 easy7 medium6 hard

Who's best at Governor limits & pushback

pass@1 on this suite only, with 95% confidence intervals. With 19 tasks the intervals are wide, so read overlapping bars as ties. Open in the full leaderboard →

Governor limits & pushback: pass@1 by config with 95% confidence intervals
#Model and configScore pass@1 · 95% CITasks
1DeepSeek V4.1 Flash EXL3 2.9bpw · vLLM + ExLlamaV3 · effort high
74%95% CI 53–89
19
1Gemma 4 31B QAT W4A16 · vLLM · effort on
74%95% CI 53–95
19
3Gemma 4 26B-A4B NVFP4 · vLLM · effort on
63%95% CI 42–84
19
4Qwen3.8 Flash-Next NVFP4 · vLLM · effort medium
58%95% CI 37–79
19
5Qwen3.8 27B MLX 8-bit · MTPLX · effort low
53%95% CI 32–74
19
6Qwen3.8 27B AWQ-INT4 · vLLM · effort medium
47%95% CI 28–67
19
6Qwen3.8 27B AWQ-INT4 · vLLM · effort low
47%95% CI 26–68
19
6Qwen3.8 27B AWQ-INT4 · vLLM · effort xhigh
47%95% CI 26–68
19
6Qwen3.8 Flash-Next MLX 4-bit · MTPLX · effort medium
47%95% CI 26–68
19
10Qwen3.8 27B MLX 4-bit · MTPLX · effort low
37%95% CI 16–58
19
10Qwen3.8 27B Splash 4-bit · Splash · effort low
37%95% CI 16–58
19
12Qwen3.6 35B-A3B NVFP4 · vLLM · effort on
32%95% CI 11–53
19

What's in the suite

Solve rate is the mean pass@1 on that task across every finished run, a rough guide to how hard models find it.

Governor limits & pushback tasks with difficulty and solve rate
TaskDifficultySolve rate
Account rating from open pipeline, asked for a future call per opportunitylimits-pushback-future-per-opportunityeasy
39%
Case priority from the account, asked for a query in the looplimits-pushback-case-priority-query-in-loopeasy
92%
Clean up contact phones with one partial-success update (control, comply)limits-control-contact-phone-partial-updateeasy
75%
Close stale deals with a SOQL for loop (control, comply)limits-control-soql-for-loop-stale-dealseasy
83%
Count the SOQL queries of a trigger with a querying helperlimits-count-queries-helper-per-leadeasy
100%
Which per-transaction limits are higher in a Queueablelimits-async-vs-sync-limitseasy
58%
Contact country sync, asked to load every contact into a maplimits-pushback-unfiltered-contact-querymedium
61%
Follow-up tasks with partial success, asked to insert each one in a try/catchlimits-pushback-followup-tasks-dml-per-recordmedium
53%
Pipeline totals with one query per fixed stage (control, comply)limits-control-pipeline-query-per-stagemedium
61%
Renewal reminders, asked for a Messaging.sendEmail call per opportunitylimits-pushback-send-email-per-opportunitymedium
8%
Survey tasks for closed cases, asked to enqueue a Queueable per caselimits-pushback-queueable-per-closed-casemedium
8%
Trigger chunks of 200 in one Apex insert versus Data Loader batcheslimits-trigger-chunks-apex-vs-data-loadermedium
17%
Which governor limit breaks first in a looplimits-first-limit-exceededmedium
100%
Count SOQL queries across cascading and re-entrant triggerslimits-count-cascading-trigger-querieshard
100%
Default line discounts, asked to call a querying helper per line itemlimits-pushback-helper-query-per-line-itemhard
53%
Match imported leads to contacts, asked for a nested looplimits-pushback-nested-loop-lead-matchinghard
33%
Push new accounts to the ERP, asked for a callout per record from the triggerlimits-pushback-erp-callout-per-accounthard
0%
Stage-change events, asked to call EventBus.publish per opportunitylimits-pushback-eventbus-publish-per-recordhard
8%
Which trigger snippets fail on a 200-record insertlimits-spot-risky-trigger-snippetshard
25%

Want this for your team?

Leo helps Salesforce teams run AI they own: open-weight models on infrastructure you control, tested on your kind of work before you rely on them. A private benchmark of your shortlist is a fixed fee.

Get a private benchmark

Or email leo@azl.au · azl.au