Framework for creating and running LLM evaluation pipelines, graders, and result analysis.

Install

mkdir -p .claude/skills/openjudge && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/11448" && unzip -o skill.zip -d .claude/skills/openjudge && rm skill.zip

Installs to .claude/skills/openjudge

Activation

This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.

Build custom LLM evaluation pipelines using the OpenJudge framework. Covers selecting and configuring graders (LLM-based, function-based, agentic), running batch evaluations with GradingRunner, combining scores with aggregators, applying evaluation strategies (voting, average), auto-generating graders from data, and analyzing results (pairwise win rates, statistics, validation metrics). Use when the user wants to evaluate LLM outputs, compare multiple models, design scoring criteria, or build an automated evaluation system.
529 chars✓ has a “when” triggerlonger than Claude Code's old 250-char listing cap (fine on current versions)
Intermediate

Key capabilities

  • Select and configure LLM-based graders
  • Run batch evaluations with GradingRunner
  • Combine scores with aggregators
  • Apply evaluation strategies like voting or averaging
  • Analyze results for pairwise win rates and statistics

How it works

The skill orchestrates LLM evaluation pipelines by configuring graders, running batch evaluations on datasets, and then analyzing the results using various strategies and aggregators.

Inputs & outputs

You give it
Dataset of queries and responses, grader configurations, LLM model configuration
You get back
GraderScore, GraderRank, GraderError, RunnerResult, statistical analysis of results

When to use openjudge

  • Evaluate LLM outputs
  • Compare multiple models
  • Build scoring rubrics

About this skill

OpenJudge Skill

Build evaluation pipelines for LLM applications using the openjudge library.

When to Use This Skill

  • User wants to evaluate LLM output quality (correctness, relevance, hallucination, etc.)
  • User wants to compare two or more models and rank them
  • User wants to design a scoring rubric and automate evaluation
  • User wants to analyze evaluation results statistically
  • User wants to build a reward model or quality filter

Sub-documents — Read When Relevant

TopicFileRead when…
Grader selection & configurationgraders.mdUser needs to pick or configure an evaluator
Batch evaluation pipelinepipeline.mdUser needs to run evaluation over a dataset
Auto-generate graders from datagenerator.mdNo rubric yet; generate from labeled examples
Analyze & compare resultsanalyzer.mdUser wants win rates, statistics, or metrics

Read the relevant sub-document before writing any code.

Install

pip install py-openjudge

Architecture Overview

Dataset (List[dict])
    │
    ▼
GradingRunner                    ← orchestrates everything
    │
    ├─► Grader A ──► EvaluationStrategy ──► _aevaluate() ──► GraderScore / GraderRank
    ├─► Grader B ──► EvaluationStrategy ──► _aevaluate() ──► GraderScore / GraderRank
    └─► Grader C ...
    │
    ├─► Aggregator (optional)    ← combine multiple grader scores into one
    │
    └─► RunnerResult             ← {grader_name: [GraderScore, ...]}
            │
            ▼
        Analyzer                 ← statistics, win rates, validation metrics

5-Minute Quick Start

Evaluate responses for correctness using a built-in grader:

import asyncio
from openjudge.models.openai_chat_model import OpenAIChatModel
from openjudge.graders.common.correctness import CorrectnessGrader
from openjudge.runner.grading_runner import GradingRunner

# 1. Configure the judge model (OpenAI-compatible endpoint)
model = OpenAIChatModel(
    model="qwen-plus",
    api_key="sk-xxx",
    base_url="https://dashscope.aliyuncs.com/compatible-mode/v1",
)

# 2. Instantiate a grader
grader = CorrectnessGrader(model=model)

# 3. Prepare dataset
dataset = [
    {
        "query": "What is the capital of France?",
        "response": "Paris is the capital of France.",
        "reference_response": "Paris.",
    },
    {
        "query": "What is 2 + 2?",
        "response": "The answer is five.",
        "reference_response": "4.",
    },
]

# 4. Run evaluation
async def main():
    runner = GradingRunner(
        grader_configs={"correctness": grader},
        max_concurrency=8,
    )
    results = await runner.arun(dataset)

    for i, result in enumerate(results["correctness"]):
        print(f"[{i}] score={result.score}  reason={result.reason}")

asyncio.run(main())

Expected output:

[0] score=5  reason=The response accurately states Paris as capital...
[1] score=1  reason=The response gives the wrong answer (five vs 4)...

Key Data Types

TypeDescription
GraderScorePointwise result: .score (float), .reason (str), .metadata (dict)
GraderRankListwise result: .rank (List[int]), .reason (str), .metadata (dict)
GraderErrorError during evaluation: .error (str), .reason (str)
RunnerResultDict[str, List[GraderResult]] — keyed by grader name

Result Handling Pattern

from openjudge.graders.schema import GraderScore, GraderRank, GraderError

for grader_name, grader_results in results.items():
    for i, result in enumerate(grader_results):
        if isinstance(result, GraderScore):
            print(f"{grader_name}[{i}]: score={result.score}")
        elif isinstance(result, GraderRank):
            print(f"{grader_name}[{i}]: rank={result.rank}")
        elif isinstance(result, GraderError):
            print(f"{grader_name}[{i}]: ERROR — {result.error}")

Model Configuration

All LLM-based graders accept either a BaseChatModel instance or a dict config:

# Option A: instance
from openjudge.models.openai_chat_model import OpenAIChatModel
model = OpenAIChatModel(model="gpt-4o", api_key="sk-...")

# Option B: dict (auto-creates OpenAIChatModel)
model_cfg = {"model": "gpt-4o", "api_key": "sk-..."}
grader = CorrectnessGrader(model=model_cfg)

# OpenAI-compatible endpoints (DashScope / local / etc.)
model = OpenAIChatModel(
    model="qwen-plus",
    api_key="sk-xxx",
    base_url="https://dashscope.aliyuncs.com/compatible-mode/v1",
)

When not to use it

  • When the user does not want to evaluate LLM outputs
  • When the user does not want to compare multiple models
  • When the user does not want to automate evaluation

Prerequisites

py-openjudge

Limitations

  • Requires an OpenAI-compatible endpoint for LLM-based graders
  • Evaluation strategies are limited to voting or average
  • Analysis focuses on pairwise win rates and statistics

How it compares

This skill provides a structured framework for building custom LLM evaluation pipelines, automating the process of grading, comparing, and analyzing model outputs, which is more systematic than manual review.

Compared to similar skills

openjudge side by side with the closest alternatives in the catalog.

SkillInstallsUpdatedSafetyDifficulty
openjudge (this skill)05moReviewIntermediate
evaluating-machine-learning-models127dReviewIntermediate
llm-evaluation62moNo flagsAdvanced
evaluating-llms-harness37moReviewAdvanced

Try saying

Example prompts that trigger this skill in your AI assistant.

You might also like

Search skills

Search the agent skills registry