evaluate
Automates the evaluation of models and systems using a two-stage smoke and full-run workflow.
Install
mkdir -p .claude/skills/evaluate && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/9722" && unzip -o skill.zip -d .claude/skills/evaluate && rm skill.zipInstalls to .claude/skills/evaluate
Activation
This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.
Run benchy evaluations against models or systems. Covers the canonical smoke→full workflow, config selection, task filtering, exit policies, and reading run_outcome.json. Use when asked to evaluate, benchmark, or run benchy against a model or system config.Key capabilities
- →Smoke test execution
- →Full-scale evaluation
- →Config selection
- →Task filtering
- →Outcome analysis
How it works
It enforces a two-stage workflow (smoke test then full run) to validate model performance against benchmarks.
Inputs & outputs
When to use evaluate
- →Running automated model benchmarks
- →Executing smoke tests on model configs
- →Validating model performance after updates
- →Analyzing evaluation outcomes and error logs
About evaluate
This skill manages the execution of benchy evaluations to ensure model performance meets requirements. It enforces a strict workflow requiring a successful smoke test before proceeding to a full-scale evaluation.
Run benchy evaluations against models or systems. Covers the canonical smoke→full workflow, config selection, task filtering, exit policies, and reading run_outcome.json. Use when asked to evaluate, benchmark, or run benchy against a model or system config.
When not to use it
- →Non-benchy evaluation frameworks
Prerequisites
Limitations
- →Requires strict adherence to the smoke test pass criteria
How it compares
It mandates a strict, automated validation workflow rather than ad-hoc testing.
Compared to similar skills
evaluate side by side with the closest alternatives in the catalog.
| Skill | Installs | Updated | Safety | Difficulty |
|---|---|---|---|---|
| evaluate (this skill) | 0 | 3mo | Review | Intermediate |
| agent-evaluation | 3 | 8mo | No flags | Advanced |
| phoenix-evals | 3 | 2mo | No flags | Advanced |
| mlops-validation | 2 | 8mo | No flags | Intermediate |
Try saying
Example prompts that trigger this skill in your AI assistant.
More by surus-lat
View all by surus-lat →You might also like
agent-evaluation
davila7
Testing and benchmarking LLM agents including behavioral testing, capability assessment, reliability metrics, and production monitoring—where even top agents achieve less than 50% on real-world benchmarks Use when: agent testing, agent evaluation, benchmark agents, agent reliability, test agent.
phoenix-evals
Arize-ai
Build and run evaluators for AI/LLM applications using Phoenix.
mlops-validation
fmind
Guide to implement rigorous validation layers including static analysis, automated testing, structured logging, and security scanning.
evaluating-machine-learning-models
jeremylongshore
Build this skill allows AI assistant to evaluate machine learning models using a comprehensive suite of metrics. it should be used when the user requests model performance analysis, validation, or testing. AI assistant can use this skill to assess model accuracy, p... Use when appropriate context detected. Trigger with relevant phrases based on skill purpose.
create-eval
HolmesGPT
This skill should be used when the user asks to "create an eval", "write an eval test", "add a new eval", "create a test case", "write a test for Holmes", or discusses LLM evaluation tests, eval fixtures, or test_case.yaml files for the HolmesGPT project.
agent-eval-api
Omniloy
Operate the Omniloy Agent Evaluator (Agent Testing Platform) API end-to-end — authenticate, resolve or create agents / personas / evaluators / test configs, launch a test run, poll it to a terminal state, and read back transcripts, scores and pass/fail. Use this whenever the user wants to test or ev