Automates the evaluation of models and systems using a two-stage smoke and full-run workflow.

Install

mkdir -p .claude/skills/evaluate && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/9722" && unzip -o skill.zip -d .claude/skills/evaluate && rm skill.zip

Installs to .claude/skills/evaluate

Activation

This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.

Run benchy evaluations against models or systems. Covers the canonical smoke→full workflow, config selection, task filtering, exit policies, and reading run_outcome.json. Use when asked to evaluate, benchmark, or run benchy against a model or system config.
257 chars✓ has a “when” triggerlonger than Claude Code's old 250-char listing cap (fine on current versions)
Intermediate

Key capabilities

  • →Smoke test execution
  • →Full-scale evaluation
  • →Config selection
  • →Task filtering
  • →Outcome analysis

How it works

It enforces a two-stage workflow (smoke test then full run) to validate model performance against benchmarks.

Inputs & outputs

You give it
Model or system configuration
You get back
Evaluation outcome report

When to use evaluate

  • →Running automated model benchmarks
  • →Executing smoke tests on model configs
  • →Validating model performance after updates
  • →Analyzing evaluation outcomes and error logs

About evaluate

This skill manages the execution of benchy evaluations to ensure model performance meets requirements. It enforces a strict workflow requiring a successful smoke test before proceeding to a full-scale evaluation.

Run benchy evaluations against models or systems. Covers the canonical smoke→full workflow, config selection, task filtering, exit policies, and reading run_outcome.json. Use when asked to evaluate, benchmark, or run benchy against a model or system config.

When not to use it

  • →Non-benchy evaluation frameworks

Prerequisites

Benchy environmentConfigured API keys

Limitations

  • →Requires strict adherence to the smoke test pass criteria

How it compares

It mandates a strict, automated validation workflow rather than ad-hoc testing.

Compared to similar skills

evaluate side by side with the closest alternatives in the catalog.

SkillInstallsUpdatedSafetyDifficulty
evaluate (this skill)03moReviewIntermediate
agent-evaluation38moNo flagsAdvanced
phoenix-evals32moNo flagsAdvanced
mlops-validation28moNo flagsIntermediate

Try saying

Example prompts that trigger this skill in your AI assistant.

You might also like

agent-evaluation

davila7

Testing and benchmarking LLM agents including behavioral testing, capability assessment, reliability metrics, and production monitoring—where even top agents achieve less than 50% on real-world benchmarks Use when: agent testing, agent evaluation, benchmark agents, agent reliability, test agent.

331

phoenix-evals

Arize-ai

Build and run evaluators for AI/LLM applications using Phoenix.

319

mlops-validation

fmind

Guide to implement rigorous validation layers including static analysis, automated testing, structured logging, and security scanning.

28

evaluating-machine-learning-models

jeremylongshore

Build this skill allows AI assistant to evaluate machine learning models using a comprehensive suite of metrics. it should be used when the user requests model performance analysis, validation, or testing. AI assistant can use this skill to assess model accuracy, p... Use when appropriate context detected. Trigger with relevant phrases based on skill purpose.

13

create-eval

HolmesGPT

This skill should be used when the user asks to "create an eval", "write an eval test", "add a new eval", "create a test case", "write a test for Holmes", or discusses LLM evaluation tests, eval fixtures, or test_case.yaml files for the HolmesGPT project.

10

agent-eval-api

Omniloy

Operate the Omniloy Agent Evaluator (Agent Testing Platform) API end-to-end — authenticate, resolve or create agents / personas / evaluators / test configs, launch a test run, poll it to a terminal state, and read back transcripts, scores and pass/fail. Use this whenever the user wants to test or ev

00

Search skills

Search the agent skills registry