openjudge
Framework for creating and running LLM evaluation pipelines, graders, and result analysis.
Install
mkdir -p .claude/skills/openjudge && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/11448" && unzip -o skill.zip -d .claude/skills/openjudge && rm skill.zipInstalls to .claude/skills/openjudge
Activation
This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.
Build custom LLM evaluation pipelines using the OpenJudge framework. Covers selecting and configuring graders (LLM-based, function-based, agentic), running batch evaluations with GradingRunner, combining scores with aggregators, applying evaluation strategies (voting, average), auto-generating graders from data, and analyzing results (pairwise win rates, statistics, validation metrics). Use when the user wants to evaluate LLM outputs, compare multiple models, design scoring criteria, or build an automated evaluation system.Key capabilities
- →Select and configure LLM-based graders
- →Run batch evaluations with GradingRunner
- →Combine scores with aggregators
- →Apply evaluation strategies like voting or averaging
- →Analyze results for pairwise win rates and statistics
How it works
The skill orchestrates LLM evaluation pipelines by configuring graders, running batch evaluations on datasets, and then analyzing the results using various strategies and aggregators.
Inputs & outputs
When to use openjudge
- →Evaluate LLM outputs
- →Compare multiple models
- →Build scoring rubrics
About openjudge
Orchestrates batch evaluations, configures LLM/function graders, and applies statistical analysis to compare outputs.
>
When not to use it
- →When the user does not want to evaluate LLM outputs
- →When the user does not want to compare multiple models
- →When the user does not want to automate evaluation
Prerequisites
Limitations
- →Requires an OpenAI-compatible endpoint for LLM-based graders
- →Evaluation strategies are limited to voting or average
- →Analysis focuses on pairwise win rates and statistics
How it compares
This skill provides a structured framework for building custom LLM evaluation pipelines, automating the process of grading, comparing, and analyzing model outputs, which is more systematic than manual review.
Compared to similar skills
openjudge side by side with the closest alternatives in the catalog.
| Skill | Installs | Updated | Safety | Difficulty |
|---|---|---|---|---|
| openjudge (this skill) | 0 | 6mo | Review | Intermediate |
| evaluating-machine-learning-models | 1 | 2mo | Review | Intermediate |
| llm-evaluation | 6 | 4mo | No flags | Advanced |
| evaluating-llms-harness | 3 | 8mo | Review | Advanced |
Try saying
Example prompts that trigger this skill in your AI assistant.
More by agentscope-ai
View all by agentscope-ai →You might also like
evaluating-machine-learning-models
jeremylongshore
Build this skill allows AI assistant to evaluate machine learning models using a comprehensive suite of metrics. it should be used when the user requests model performance analysis, validation, or testing. AI assistant can use this skill to assess model accuracy, p... Use when appropriate context detected. Trigger with relevant phrases based on skill purpose.
llm-evaluation
wshobson
Implement comprehensive evaluation strategies for LLM applications using automated metrics, human feedback, and benchmarking. Use when testing LLM performance, measuring AI application quality, or establishing evaluation frameworks.
evaluating-llms-harness
davila7
Evaluates LLMs across 60+ academic benchmarks (MMLU, HumanEval, GSM8K, TruthfulQA, HellaSwag). Use when benchmarking model quality, comparing models, reporting academic results, or tracking training progress. Industry standard used by EleutherAI, HuggingFace, and major labs. Supports HuggingFace, vLLM, APIs.
llm-evaluation
H4D3ZS
Implement comprehensive evaluation strategies for LLM applications using automated metrics, human feedback, and benchmarking. Use when testing LLM performance, measuring AI application quality, or establishing evaluation frameworks.
Plate Evaluation
ShArAvaNPai
Runs current model against validation set and returns JSON metrics with automated recommendations
model-change
connorkitchings
Use for prediction logic, feature usage, backtests, or metrics changes.