Suites / ci
CI/CD
Salesforce CI/CD engineering: GitHub Actions workflows and pipeline scripts that authenticate with JWT or SFDX auth URLs from secrets, validate pull requests, quick deploy validated releases, run delta deployments with sfdx-git-delta, manage scratch orgs per pull request, build and promote unlocked packages, gate builds on Code Analyzer and run LWC Jest, plus CI/CD judgement questions. Every sf command in an answer is validated against the pinned sf CLI manifest; workflow properties (triggers, job ordering, environments, cleanup, secrets) are checked deterministically, including simulated events. No org is needed.
How it's graded
GitHub Actions workflows are checked for structure and secrets hygiene, and every sf command in them is validated against the pinned sf command manifest.
Difficulty mix
Leaderboard
Who's best at CI/CD
pass@1 on this suite only, with 95% confidence intervals. With 17 tasks the intervals are wide, so read overlapping bars as ties. Open in the full leaderboard →
| # | Model and config | Score pass@1 · 95% CI | Tasks |
|---|---|---|---|
| 1 | DeepSeek V4.1 Flash EXL3 2.9bpw · vLLM + ExLlamaV3 · effort high | 71%95% CI 47–88 | 17 |
| 2 | Gemma 4 31B QAT W4A16 · vLLM · effort on | 59%95% CI 35–82 | 17 |
| 2 | Qwen3.8 Flash-Next NVFP4 · vLLM · effort medium | 59%95% CI 35–82 | 17 |
| 4 | Qwen3.8 Flash-Next MLX 4-bit · MTPLX · effort medium | 47%95% CI 24–71 | 17 |
| 5 | Qwen3.8 27B AWQ-INT4 · vLLM · effort low | 41%95% CI 18–65 | 17 |
| 5 | Qwen3.8 27B AWQ-INT4 · vLLM · effort xhigh | 41%95% CI 18–65 | 17 |
| 7 | Gemma 4 26B-A4B NVFP4 · vLLM · effort on | 35%95% CI 12–59 | 17 |
| 7 | Qwen3.8 27B AWQ-INT4 · vLLM · effort medium | 35%95% CI 12–59 | 17 |
| 7 | Qwen3.8 27B Splash 4-bit · Splash · effort low | 35%95% CI 12–59 | 17 |
| 10 | Qwen3.6 35B-A3B NVFP4 · vLLM · effort on | 29%95% CI 12–53 | 17 |
| 10 | Qwen3.8 27B MLX 8-bit · MTPLX · effort low | 29%95% CI 12–53 | 17 |
| 10 | Qwen3.8 27B MLX 4-bit · MTPLX · effort low | 29%95% CI 12–53 | 17 |
| – | GLM-5.3 Flash partialEXL3 4.0bpw · vLLM + ExLlamaV3 · effort highRun so far: 14 of 15 suites · 219 of 272 tasks graded | 41%95% CI 18–65 | 17 |
Tasks
What's in the suite
Solve rate is the mean pass@1 on that task across every finished run, a rough guide to how hard models find it.
| Task | Difficulty | Solve rate |
|---|---|---|
| Deploy to a UAT sandbox with an SFDX auth URLci-uat-deploy-sfdx-url | easy | 33% |
| LWC lint and Jest job with cached npm dependenciesci-lwc-jest-eslint | easy | 100% |
| Nightly Apex regression run in a full sandboxci-nightly-apex-tests | easy | 0% |
| Pre or post destructive changes for a referenced fieldci-choice-destructive-order | easy | 100% |
| Test level for a production deploy with managed packagesci-choice-test-level-prod | easy | 100% |
| Block pull requests on high-severity Code Analyzer violationsci-code-analyzer-gate | medium | 0% |
| Fix a production deploy that lost its environment secretsci-environment-secrets-fix | medium | 100% |
| Scratch orgs versus sandboxes as per-pull-request CI orgsci-choice-scratch-vs-sandbox | medium | 75% |
| Script a disposable review scratch orgci-review-org-script | medium | 8% |
| Source tracking behaviour in CI jobsci-choice-source-tracking-ci | medium | 42% |
| Validate pull requests against production with JWT authci-pr-validate-jwt | medium | 8% |
| Why a quick deploy finds no recent job IDci-choice-quick-deploy-cache | medium | 100% |
| Build and promote an unlocked package version on release tagsci-unlocked-package-release | hard | 8% |
| Delta validation with sfdx-git-delta and destructive changesci-delta-deploy-sgd | hard | 0% |
| One workflow that routes validations and deploys by branchci-branch-routed-validation | hard | 25% |
| Throwaway scratch org per pull request with guaranteed cleanupci-scratch-org-per-pr | hard | 8% |
| Validate on merge, then quick deploy after approvalci-quick-deploy-on-merge | hard | 17% |
Work with Leo
Want this for your team?
Leo helps Salesforce teams run AI they own: open-weight models on infrastructure you control, tested on your kind of work before you rely on them. A private benchmark of your shortlist is a fixed fee.
Get a private benchmarkOr email leo@azl.au · azl.au