agentic-eval Skill
在现有代码库中规划、实现或评审智能体评估工作,并提供兼容性、安全性和验证控制。仅在用户明确要求智能体评估工作时使用。
安装方式:把技能目录放入 ~/.claude/skills/(Claude Code)或在 claude.ai 设置中启用;也可复制右侧安装命令一键添加。
技能指令原文(SKILL.md)
Agentic Eval
Use this Skill to produce a bounded, verifiable Agentic Eval outcome. Preserve the user's chosen stack, source material, and authorization boundaries.
Read the SandBase API map only when the task genuinely needs an external data source or generative model.
Workflow
- Inspect the available files, runtime, versions, inputs, and existing conventions before deciding what to change.
- Restate the requested outcome, constraints, acceptance checks, and any assumption that could change the result.
- Produce the smallest complete implementation, analysis, or artifact that satisfies those checks.
- Verify the real output with appropriate tests, previews, calculations, or source comparison; do not infer success from file creation alone.
- Return the deliverable, evidence of validation, material assumptions, and unresolved limitations.
Quality gates
- Inspect repository instructions, current versions, tests, and user changes before proposing an implementation.
- Make the smallest compatible change, preserve public contracts unless change is requested, and avoid new dependencies without a concrete benefit.
- Run focused tests plus the relevant build or static checks, then report observed results and any unverified paths.
- Define measurable acceptance criteria before evaluation, retain counterexamples and uncertainty, and do not collapse important failures into an average score.
SandBase boundary
Keep the core Agentic Eval work local. Use SandBase only for current external documentation, repository evidence, or explicitly requested model inference.
- Call
sandbase_discoverwith a short capability query. - Call
sandbase_inspectfor viable candidates and compare the live schema, coverage, limits, output, execution mode, and price. - Prefer a dedicated tool or API the user already has. Send only the minimum necessary data.
- Before any paid call, show the endpoint, important arguments, current unit price, call count, and total estimate or uncertainty, then obtain confirmation.
- Use
sandbase_accountbefore an approved multi-call batch and callsandbase_runonly with current schema-defined arguments. - Poll asynchronous work with
sandbase_run_getusing the same run ID; never resubmit merely because it is pending. - Use
sandbase_runsonly to recover status or reconcile observed cost.
If SandBase is unavailable, continue with local work and authorized sources when possible. Do not silently switch providers, fabricate external results, or claim a generation or retrieval succeeded.
Handoff
Provide the completed artifact or findings, concise reproduction steps, checks actually run, source or asset provenance, SandBase endpoint and run IDs when used, observed cost when available, and any follow-up that still requires user action.