Suites / fflib
fflib Enterprise Patterns
Apex Enterprise Patterns with the open-source fflib libraries (fflib-apex-common and fflib-apex-mocks): Selector, Domain, Service and Unit of Work layers, the Application factory class, and unit tests with ApexMocks. Code is deployed to a scratch org with the pinned fflib commits installed and must pass hidden Apex tests; test-writing tasks are graded by mutation testing (the model's tests must pass against the correct code and fail against seeded bugs).
How it's graded
Code is deployed as a check-only validation to a scratch org with the pinned fflib libraries installed and must pass hidden Apex tests. ApexMocks test-writing tasks are graded by mutation testing: the model’s tests must pass against the correct code and fail against seeded bugs.
Difficulty mix
Leaderboard
Who's best at fflib Enterprise Patterns
pass@1 on this suite only, with 95% confidence intervals. With 18 tasks the intervals are wide, so read overlapping bars as ties. Open in the full leaderboard →
| # | Model and config | Score pass@1 · 95% CI | Tasks |
|---|---|---|---|
| 1 | DeepSeek V4.1 Flash EXL3 2.9bpw · vLLM + ExLlamaV3 · effort high | 44%95% CI 22–67 | 18 |
| 2 | Qwen3.6 35B-A3B NVFP4 · vLLM · effort on | 11%95% CI 0–28 | 18 |
| 2 | Qwen3.8 Flash-Next MLX 4-bit · MTPLX · effort medium | 11%95% CI 0–28 | 18 |
| 4 | Gemma 4 31B QAT W4A16 · vLLM · effort on | 6%95% CI 0–17 | 18 |
| 4 | Qwen3.8 27B MLX 4-bit · MTPLX · effort low | 6%95% CI 0–17 | 18 |
| 4 | Qwen3.8 Flash-Next NVFP4 · vLLM · effort medium | 6%95% CI 0–17 | 18 |
| 7 | Gemma 4 26B-A4B NVFP4 · vLLM · effort on | 0%95% CI 0–0 | 18 |
| 7 | Qwen3.8 27B MLX 8-bit · MTPLX · effort low | 0%95% CI 0–0 | 18 |
| 7 | Qwen3.8 27B AWQ-INT4 · vLLM · effort medium | 0%95% CI 0–0 | 18 |
| 7 | Qwen3.8 27B Splash 4-bit · Splash · effort low | 0%95% CI 0–0 | 18 |
| 7 | Qwen3.8 27B AWQ-INT4 · vLLM · effort low | 0%95% CI 0–0 | 18 |
| 7 | Qwen3.8 27B AWQ-INT4 · vLLM · effort xhigh | 0%95% CI 0–0 | 18 |
| – | GLM-5.3 Flash partialEXL3 4.0bpw · vLLM + ExLlamaV3 · effort highRun so far: 14 of 15 suites · 219 of 272 tasks graded | 22%95% CI 0–56 | 9 |
Tasks
What's in the suite
Solve rate is the mean pass@1 on that task across every finished run, a rough guide to how hard models find it.
| Task | Difficulty | Solve rate |
|---|---|---|
| ApexMocks unit test with a stubbed selectorfflib-apexmocks-stub-selector | easy | 8% |
| Contact selector with fflib_SObjectSelectorfflib-selector-contacts | easy | 8% |
| Create accounts and contacts in one unit of workfflib-uow-account-contacts | easy | 8% |
| Opportunity domain class with insert defaultsfflib-domain-opportunity-defaults | easy | 8% |
| Static service facade over the Application service factoryfflib-service-facade | easy | 42% |
| ApexMocks test verifying unit of work interactionsfflib-apexmocks-verify-uow | medium | 0% |
| Application class with the four fflib factoriesfflib-application-factories | medium | 0% |
| Custom selector query with its own ordering and limitfflib-selector-query-factory | medium | 0% |
| Domain trigger logic delegating to a service once per triggerfflib-domain-calls-service | medium | 8% |
| Domain unit test with fflib's mock trigger databasefflib-domain-unit-test-mock-db | medium | 8% |
| Insert and update validation in an fflib domain classfflib-domain-validation | medium | 0% |
| Mockable service using the Application unit of work and selectorfflib-service-case-escalation | medium | 33% |
| ApexMocks test for a service's dependency failuresfflib-apexmocks-exception-paths | hard | 0% |
| ApexMocks test for records created and related through the unit of workfflib-apexmocks-argument-captor | hard | 0% |
| ApexMocks test that mocks a domain and verifies call orderfflib-apexmocks-mock-domain | hard | 0% |
| Mockable domain on the fflib_SObjects structurefflib-domain-sobjects-structure | hard | 0% |
| Publish platform events only after a successful unit of work commitfflib-uow-platform-events | hard | 0% |
| User-mode selector with a field set and a child subselectfflib-selector-user-mode-fieldset | hard | 0% |
Work with Leo
Want this for your team?
Leo helps Salesforce teams run AI they own: open-weight models on infrastructure you control, tested on your kind of work before you rely on them. A private benchmark of your shortlist is a fixed fee.
Get a private benchmarkOr email leo@azl.au · azl.au