Provides strategies to evaluate, benchmark, and test LLM performance systematically.

Install

mkdir -p .claude/skills/llm-evaluation-h4d3zs && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/13370" && unzip -o skill.zip -d .claude/skills/llm-evaluation-h4d3zs && rm skill.zip

Installs to .claude/skills/llm-evaluation-h4d3zs

Activation

This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.

Implement comprehensive evaluation strategies for LLM applications using automated metrics, human feedback, and benchmarking. Use when testing LLM performance, measuring AI application quality, or establishing evaluation frameworks.
232 chars✓ has a “when” trigger
Intermediate

Key capabilities

  • Measure LLM application performance systematically
  • Compare different models or prompts
  • Detect performance regressions before deployment
  • Validate improvements from prompt changes
  • Establish baselines and track progress over time
  • Debug unexpected model behavior

How it works

This skill provides strategies for evaluating LLM applications using automated metrics, human feedback, and benchmarking to assess performance and quality.

Inputs & outputs

You give it
LLM application performance data, model outputs, or prompt variations
You get back
Evaluation metrics, performance comparisons, or regression detection reports

When to use llm-evaluation

  • Benchmarking LLM prompts
  • Monitoring AI model quality
  • Testing for performance regressions

About this skill

LLM Evaluation

Master comprehensive evaluation strategies for LLM applications, from automated metrics to human evaluation and A/B testing.

Do not use this skill when

  • The task is unrelated to llm evaluation
  • You need a different domain or tool outside this scope

Instructions

  • Clarify goals, constraints, and required inputs.
  • Apply relevant best practices and validate outcomes.
  • Provide actionable steps and verification.
  • If detailed examples are required, open resources/implementation-playbook.md.

Use this skill when

  • Measuring LLM application performance systematically
  • Comparing different models or prompts
  • Detecting performance regressions before deployment
  • Validating improvements from prompt changes
  • Building confidence in production systems
  • Establishing baselines and tracking progress over time
  • Debugging unexpected model behavior

Core Evaluation Types

🧠 Knowledge Modules (Fractal Skills)

1. 1. Automated Metrics

2. 2. Human Evaluation

3. 3. LLM-as-Judge

4. BLEU Score

5. ROUGE Score

6. BERTScore

7. Custom Metrics

8. Single Output Evaluation

9. Pairwise Comparison

10. Annotation Guidelines

11. Inter-Rater Agreement

12. Statistical Testing Framework

13. Regression Detection

14. Running Benchmarks

Limitations

  • The task must be related to LLM evaluation
  • The skill does not cover domains or tools outside its scope

How it compares

This skill offers structured evaluation types and sub-skills for LLM applications, unlike generic performance testing.

Compared to similar skills

llm-evaluation side by side with the closest alternatives in the catalog.

SkillInstallsUpdatedSafetyDifficulty
llm-evaluation (this skill)03moNo flagsIntermediate
Plate Evaluation05moReviewIntermediate
model-change04moReviewIntermediate
llm-evaluation62moNo flagsAdvanced

Try saying

Example prompts that trigger this skill in your AI assistant.

email-sequence

H4D3ZS

When the user wants to create or optimize an email sequence, drip campaign, automated email flow, or lifecycle email program. Also use when the user mentions "email sequence," "drip campaign," "nurture sequence," "onboarding emails," "welcome sequence," "re-engagement emails," "email automation," or

00

agent-code-guide

H4D3ZS

Master guide for using Agent Code effectively. Includes configuration templates, prompting strategies "Thinking" keywords, debugging techniques, and best practices for interacting with the agent.

00

conductor-implement

H4D3ZS

Execute tasks from a track's implementation plan following TDD workflow

00

interactive-portfolio

H4D3ZS

Expert in building portfolios that actually land jobs and clients - not just showing work, but creating memorable experiences. Covers developer portfolios, designer portfolios, creative portfolios, and portfolios that convert visitors into opportunities. Use when: portfolio, personal website, showca

00

Broken Authentication Testing

H4D3ZS

This skill should be used when the user asks to "test for broken authentication vulnerabilities", "assess session management security", "perform credential stuffing tests", "evaluate password policies", "test for session fixation", or "identify authentication bypass flaws". It provides comprehensive

00

Privilege Escalation Methods

H4D3ZS

This skill should be used when the user asks to "escalate privileges", "get root access", "become administrator", "privesc techniques", "abuse sudo", "exploit SUID binaries", "Kerberoasting", "pass-the-ticket", "token impersonation", or needs guidance on post-exploitation privilege escalation for Li

00

Search skills

Search the agent skills registry