Suites / taf
Trigger Actions Framework
Real work in orgs that use the open-source Trigger Actions Framework (TAF) by Mitch Spano: writing trigger action classes, registering them with custom metadata, ordering, bypass mechanisms, entry criteria, recursion control, DML finalizers, flow actions, migrating legacy triggers and testing actions. Each answer is deployed (check-only) with hidden Apex tests to a scratch org that has TAF 0.3.4 installed; the tests perform real DML and assert that the framework ran the model's actions as specified.
How it's graded
Each answer is deployed as a check-only validation, with hidden Apex tests, to a scratch org with the Trigger Actions Framework installed. The tests do real DML and assert that the framework ran the model’s actions as specified: ordering, bypasses, entry criteria and recursion control.
Difficulty mix
Leaderboard
Who's best at Trigger Actions Framework
pass@1 on this suite only, with 95% confidence intervals. With 18 tasks the intervals are wide, so read overlapping bars as ties. Open in the full leaderboard →
| # | Model and config | Score pass@1 · 95% CI | Tasks |
|---|---|---|---|
| 1 | DeepSeek V4.1 Flash EXL3 2.9bpw · vLLM + ExLlamaV3 · effort high | 11%95% CI 0–28 | 18 |
| 2 | Gemma 4 26B-A4B NVFP4 · vLLM · effort on | 6%95% CI 0–17 | 18 |
| 2 | Qwen3.8 Flash-Next MLX 4-bit · MTPLX · effort medium | 6%95% CI 0–17 | 18 |
| 4 | Gemma 4 31B QAT W4A16 · vLLM · effort on | 0%95% CI 0–0 | 18 |
| 4 | Qwen3.6 35B-A3B NVFP4 · vLLM · effort on | 0%95% CI 0–0 | 18 |
| 4 | Qwen3.8 27B MLX 8-bit · MTPLX · effort low | 0%95% CI 0–0 | 18 |
| 4 | Qwen3.8 27B MLX 4-bit · MTPLX · effort low | 0%95% CI 0–0 | 18 |
| 4 | Qwen3.8 27B AWQ-INT4 · vLLM · effort medium | 0%95% CI 0–0 | 18 |
| 4 | Qwen3.8 27B Splash 4-bit · Splash · effort low | 0%95% CI 0–0 | 18 |
| 4 | Qwen3.8 27B AWQ-INT4 · vLLM · effort low | 0%95% CI 0–0 | 18 |
| 4 | Qwen3.8 27B AWQ-INT4 · vLLM · effort xhigh | 0%95% CI 0–0 | 18 |
| 4 | Qwen3.8 Flash-Next NVFP4 · vLLM · effort medium | 0%95% CI 0–0 | 18 |
| – | GLM-5.3 Flash partialEXL3 4.0bpw · vLLM + ExLlamaV3 · effort highRun so far: 14 of 15 suites · 219 of 272 tasks graded | 15%95% CI 0–38 | 13 |
Tasks
What's in the suite
Solve rate is the mean pass@1 on that task across every finished run, a rough guide to how hard models find it.
| Task | Difficulty | Solve rate |
|---|---|---|
| Enable the framework on Opportunitytaf-enable-opportunity | easy | 0% |
| Normalise new leads in a before-insert actiontaf-lead-normalize-before-insert | easy | 0% |
| Release a new action to pilot users onlytaf-required-permission-pilot | easy | 0% |
| Run an admin's flow as an ordered trigger actiontaf-register-flow-action | easy | 0% |
| Skip one action during a contact importtaf-bypass-action-in-import | easy | 0% |
| Switch off a misbehaving action without a code changetaf-disable-action-hotfix | easy | 0% |
| Audit amount changes once despite re-entrant updatestaf-recursion-amount-audit | medium | 0% |
| Bypass all Account automation for the integration usertaf-bypass-permission-integration | medium | 8% |
| Bypass one object's actions inside a service, safelytaf-bypass-object-ownership-service | medium | 0% |
| Create onboarding tasks from an after-insert actiontaf-after-insert-onboarding-task | medium | 0% |
| DML-less unit tests for a trigger actiontaf-dml-less-action-test | medium | 25% |
| Filter an action with an entry criteria formulataf-entry-criteria-formula | medium | 0% |
| One validation action registered for insert and updatetaf-duplicate-email-two-contexts | medium | 0% |
| Build a TAF flow action and bypass only that flowtaf-flow-action-authoring | hard | 0% |
| Migrate a legacy fat Case trigger to trigger actionstaf-migrate-legacy-case-trigger | hard | 0% |
| One async sync job per DML operation with a DML finalizertaf-dml-finalizer-erp-sync | hard | 0% |
| Roll up open cases to Account across all after contextstaf-open-case-rollup | hard | 0% |
| Share one parent query across trigger actionstaf-shared-query-singleton | hard | 0% |
Work with Leo
Want this for your team?
Leo helps Salesforce teams run AI they own: open-weight models on infrastructure you control, tested on your kind of work before you rely on them. A private benchmark of your shortlist is a fixed fee.
Get a private benchmarkOr email leo@azl.au · azl.au