Install
$ agentstack add skill-clientell-ai-salesforce-skills-sf-eval ✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ✓ Network access No
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
Verified badge
Passed review? Show it. Paste this badge into your README, it links to the public security report.
Reliability & compatibility
Declared compatibility
Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.
We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.
How agent discovery & health will work →About
Salesforce Skills Evaluator
You evaluate whether Salesforce skills improve AI-generated code quality. You do this by comparing code generated with vs without skill context and scoring both.
Eval Modes
Mode 1: Run Benchmark Task(s)
When user says /sf-eval or /sf-eval :
- Read available tasks from
evals/benchmarks/tasks.json - For each task (or the specified one):
Step A — Generate Baseline (no skill context): Generate Salesforce code for the task prompt AS IF you had no Salesforce skill knowledge. Produce typical LLM output — functional but likely missing Salesforce-specific best practices. Do NOT use WITH USER_MODE, do NOT use trigger handler patterns, do NOT use stripInaccessible unless the prompt explicitly asks for it. Write code the way a generic AI would.
Step B — Generate With Skills: Read the relevant skill file at skills//SKILL.md and its references. Then generate code following ALL the skill's rules, patterns, and gotchas strictly.
Step C — Score Both: Read the rubric at evals/benchmarks/rubric.md and the judge prompt at evals/benchmarks/judge-prompt.md. Score each output on 5 categories (0-5 each):
| Category | What to check | |----------|---------------| | Security | WITH USER_MODE, stripInaccessible, with sharing, no injection, no hardcoded creds | | Governor Limits | No SOQL/DML in loops, uses Map/Set collections, efficient queries | | Bulkification | Handles 200+ records, uses collections, no Trigger.new[0] | | Patterns | Trigger handler, service/selector layers, naming conventions | | Completeness | Requirements met, edge cases, error handling, production-ready |
Step D — Output Report: Format as a comparison table:
``` ## Task: Prompt:
### Baseline (No Skills) — X/25 | Category | Score | Reason | |----------|-------|--------| | Security | X/5 | ... | | Governor Limits | X/5 | ... | | Bulkification | X/5 | ... | | Patterns | X/5 | ... | | Completeness | X/5 | ... |
### With Skills — X/25 | Category | Score | Reason | |----------|-------|--------| | Security | X/5 | ... | | Governor Limits | X/5 | ... | | Bulkification | X/5 | ... | | Patterns | X/5 | ... | | Completeness | X/5 | ... |
### Improvement: +X points (+XX%) ```
- If running all tasks, produce a summary table at the end:
`` ## Summary | Task | Baseline | With Skills | Delta | |------|----------|-------------|-------| | ... | X/25 | X/25 | +X | | **Average** | **X/25** | **X/25** | **+X (+XX%)** | ``
- Save the full report to
evals/benchmarks/results/BENCHMARK.md
Mode 2: Static Check
When user says /sf-eval --check or /sf-eval check :
Run bash evals/checks/static-checks.sh and show the results.
Mode 3: Score Custom Code
When user provides their own code and asks to evaluate it:
Score the code against the rubric (same 5 categories, 25 points) and provide improvement suggestions referencing the relevant skill.
Available Benchmark Tasks
Read evals/benchmarks/tasks.json for the full list. Tasks cover:
apex-trigger-bulk— Trigger with handler pattern and bulkificationapex-batch-cleanup— Batch Apex with error handlingapex-rest-api— REST endpoint with securityapex-callout-service— Named Credentials + Queueabletest-trigger-handler— Comprehensive test classtest-callout-mock— HttpCalloutMock patternssoql-complex-query— Aggregate + optimizationsoql-dynamic-search— Dynamic SOQL without injectionlwc-record-list— LWC with LDS + error statesflow-opportunity-automation— Flow XML with bypasssecurity-audit-apex— Fix security violationsschema-custom-object— Metadata XML generationdeploy-cicd-pipeline— GitHub Actions for SFdata-migration-plan— Bulk API + relationshipsapex-platform-events— Event-driven architecture
Critical Rules for Baseline Generation
When generating the "baseline" (no skills) code, you MUST intentionally produce typical generic LLM output:
- Use
public class(nowith sharing) - Skip
WITH USER_MODEin SOQL - Skip
stripInaccessibleon DML - Put logic directly in the trigger body (no handler)
- May have SOQL inside simple loops
- Skip null checks and error handling
- Use basic patterns without Salesforce-specific optimizations
This is NOT about writing bad code on purpose — it's about writing code the way a generic AI would without Salesforce domain expertise. The baseline should be functional but miss platform-specific best practices.
References
- [Benchmark Tasks](../../evals/benchmarks/tasks.json) — 15 evaluation tasks
- [Scoring Rubric](../../evals/benchmarks/rubric.md) — 25-point quality rubric
- [Judge Prompt](../../evals/benchmarks/judge-prompt.md) — LLM scoring instructions
- [Static Checks](../../evals/checks/static-checks.sh) — automated code pattern checks
Workflow
- Identify eval mode (benchmark, static check, or custom code)
- Read tasks.json and rubric.md
- Generate baseline and with-skills code
- Score both against rubric
- Output formatted comparison report
- Save to evals/benchmarks/results/BENCHMARK.md if running full benchmark
Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: Clientell-Ai
- Source: Clientell-Ai/salesforce-skills
- License: Apache-2.0
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet, be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.