SOQL

Text-to-SOQL on a realistic B2B sales and service org: filters, relationship queries in both directions, semi- and anti-joins, aggregates, date and fiscal functions, polymorphic relationships, multi-select picklists and null handling. Each answer is executed in a seeded scratch org (orgs/base) and its result set must equal the gold query's (execution accuracy).

20
tasks
Execution
grading
85%
best pass@1 · DeepSeek V4.1 Flash
13
configs with results (1 partial)

How it's graded

Each query runs against a scratch org seeded with fixed data, and its result set must equal the gold query’s (execution accuracy, as in BIRD and Spider). A query that errors or returns different rows fails.

Difficulty mix

6 easy8 medium6 hard

Who's best at SOQL

pass@1 on this suite only, with 95% confidence intervals. With 20 tasks the intervals are wide, so read overlapping bars as ties. Open in the full leaderboard →

SOQL: pass@1 by config with 95% confidence intervals
#Model and configScore pass@1 · 95% CITasks
1DeepSeek V4.1 Flash EXL3 2.9bpw · vLLM + ExLlamaV3 · effort high
85%95% CI 70–100
20
2Qwen3.8 Flash-Next NVFP4 · vLLM · effort medium
75%95% CI 55–90
20
3Gemma 4 26B-A4B NVFP4 · vLLM · effort on
70%95% CI 50–90
20
3Qwen3.8 Flash-Next MLX 4-bit · MTPLX · effort medium
70%95% CI 50–90
20
5Gemma 4 31B QAT W4A16 · vLLM · effort on
65%95% CI 45–85
20
6Qwen3.6 35B-A3B NVFP4 · vLLM · effort on
55%95% CI 35–75
20
6Qwen3.8 27B MLX 8-bit · MTPLX · effort low
55%95% CI 35–75
20
8Qwen3.8 27B AWQ-INT4 · vLLM · effort medium
48%95% CI 30–68
20
9Qwen3.8 27B Splash 4-bit · Splash · effort low
45%95% CI 25–65
20
10Qwen3.8 27B AWQ-INT4 · vLLM · effort low
40%95% CI 20–60
20
11Qwen3.8 27B MLX 4-bit · MTPLX · effort low
35%95% CI 15–55
20
11Qwen3.8 27B AWQ-INT4 · vLLM · effort xhigh
35%95% CI 15–55
20
–GLM-5.3 Flash partialEXL3 4.0bpw · vLLM + ExLlamaV3 · effort highRun so far: 14 of 15 suites · 219 of 272 tasks graded
85%95% CI 70–100
20

What's in the suite

Solve rate is the mean pass@1 on that task across every finished run, a rough guide to how hard models find it.

SOQL tasks with difficulty and solve rate
TaskDifficultySolve rate
Accounts missing billing country or industrysoql-accounts-missing-country-or-industryeasy
92%
Count open German web leads for a dashboard tilesoql-open-german-web-leads-counteasy
33%
Energy-sector wins in calendar 2025soql-energy-wins-2025easy
81%
Financial-services accounts in Germany and the UKsoql-banking-insurance-de-ukeasy
100%
In-transit shipments with their account namesoql-in-transit-shipmentseasy
69%
Managers at healthcare accountssoql-healthcare-manager-contactseasy
92%
2025 closed-won business by industrysoql-won-2025-by-industrymedium
42%
Accounts with open high-priority casessoql-accounts-open-high-casesmedium
92%
Customers with no wins in 2025soql-customers-without-2025-winsmedium
61%
Five largest customers by annual revenuesoql-top-customers-by-revenuemedium
100%
German customers with their open shipments nestedsoql-german-customers-open-shipmentsmedium
0%
Open tasks related to opportunitiessoql-open-tasks-on-opportunitiesmedium
92%
Products whose code starts with SVC_soql-service-product-codesmedium
44%
Responders to an event campaignsoql-summit-respondersmedium
36%
Activity log with type-specific related-record fieldssoql-march-2025-activity-typeofhard
25%
Bookings per fiscal quarter with a non-calendar fiscal yearsoql-fy2026-bookings-by-quarterhard
25%
Contacts anywhere in an account hierarchysoql-globex-group-contactshard
58%
Hardware deals without services attachedsoql-hardware-without-services-oppshard
81%
Open shipments needing special handling (multi-select picklist)soql-open-shipments-handlinghard
8%
Parent accounts of German subsidiariessoql-parents-of-german-accountshard
0%

Want this for your team?

Leo helps Salesforce teams run AI they own: open-weight models on infrastructure you control, tested on your kind of work before you rely on them. A private benchmark of your shortlist is a fixed fee.

Get a private benchmark

Or email leo@azl.au · azl.au