evaluate-environments
Manage and analyze evaluations for AI agent tasksets using the prime CLI for benchmarking and debugging.
Install
mkdir -p .claude/skills/evaluate-environments && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/4589" && unzip -o skill.zip -d .claude/skills/evaluate-environments && rm skill.zipInstalls to .claude/skills/evaluate-environments
Activation
This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.
Run and evaluate verifiers tasksets. Set up the necessary config files and observe the runs and their results.Key capabilities
- →Validate evaluation configurations without model calls
- →Run model-free gold validation for tasksets
- →Perform small runs to verify evaluation correctness
- →Inspect successful, zero-reward, and errored traces
- →Configure taskset, harness, and runtime settings
- →Set sampling parameters for model evaluations
How it works
The skill uses the `prime eval` CLI to run evaluations, allowing dry-runs, small sample testing, and full runs with configurable taskset, harness, runtime, and sampling parameters.
Inputs & outputs
When to use evaluate-environments
- →Execute benchmark sweeps for AI tasksets
- →Run dry-runs to validate evaluation configurations
- →Analyze failure traces and sample-level outputs
- →Compare performance across different models or harnesses
About this skill
Evaluate Tasksets
Goal
Set up an evaluation for a taskset in the correct way to reproduce results from others or evaluate a model and harness combination on a given taskset.
Canonical path
Use the eval entrypoint
uv run vf-eval <MY_ENV>
Core workflow
- Resolve and validate config without model calls:
uv run vf-eval <MY_ENV> --dry-run
- Run model-free validation. Each task gets two checks in independent runtimes: gold (
setup, thenvalidateapplies and checks the reference answer; unchecked when the task has novalidate) and noop (setup, thenfinalizeand the task's full scoring, configured judges included, on the untouched task; invalid when its reward already reaches 1.0, unchecked when it has no reward).--only-gold/--only-nooprun one;--only-setuponly checks thatsetupcompletes:
uv run vf-validate <MY_ENV> --runtime.type subprocess
- Do a small run to see whether it works correctly:
uv run vf-eval <MY_ENV> -m deepseek/deepseek-v4-flash -n 3 -r 1
- Inspect successful, zero-reward, and errored traces.
- Scale only after task loading, harness capability, runtime lifecycle, and scoring are correct.
When the user requests a full run, do not restrict the number of tasks. Ask for the appropriate harness to use (if not specified)
IDs and plugin resolution
A plugin id names an installed package (e.g. my-taskset); verifiers imports it and never installs anything itself.
The leading ID is shorthand for --env.taskset.id. A harness belongs to an agent — --env.agent.harness.* on the single-agent env, --env.<agent>.harness.* on a multi-agent one (there is no run-level --harness.*):
uv run vf-eval my-task-v1 --env.agent.harness.id codex --env.agent.runtime.type prime
The env — the control flow between agents — owns the whole [env] block. Empty --env.id
keeps the taskset's own story (its exported Env subclass, else the single-agent
env); --env.id pairs a reusable env with any taskset, its knobs typed under --env.*:
uv run vf-eval my-task-v1 --env.id best-of-n --env.n 8 # pass@k / rejection sampling
uv run vf-eval my-task-v1 --env.id agentic-judge \
--env.judge.runtime.type docker # a judge agent verifies each attempt in a sandbox
Disabling tools
Almost every harness comes with a disabled_tools list, which can be used to disable one or multiple tools:
[env.agent.harness]
disabled_tools = ["shell_tool"]
The names of these tools are set by the respective harness. Research the relevant first party documentation for the given harness for the relevant name(s). Some harnesses do not offer support to disable tools.
Config discovery
The CLI help is generated from the current config classes. Include the taskset and env ids you plan to use before --help so their concrete config fields are loaded:
uv run vf-eval my-task-v1 \
--env.id best-of-n \
--help
For implementation details and defaults, start at verifiers/v1/configs/cli/eval.py and follow its fields into verifiers/v1/configs/. Client configs live in verifiers/v1/configs/client.py, sampling in verifiers/v1/types.py, and runtime- and harness-specific configs next to their implementations in verifiers/v1/runtimes/ and verifiers/v1/harnesses/. Custom taskset and env config fields live next to those implementations.
Typed taskset overrides
Taskset settings:
uv run vf-eval my-task-v1 --env.taskset.split test --env.taskset.difficulty hard
Harness and runtime settings:
uv run vf-eval my-task-v1 \
--env.agent.harness.id rlm \
--env.agent.runtime.type docker \
--env.agent.runtime.cpu 4 \
--env.agent.runtime.memory 8
Sampling:
uv run vf-eval my-task-v1 \
--sampling.temperature 0.7 \
--sampling.top-p 0.95 \
--sampling.max-tokens 2048 \
--sampling.reasoning-effort medium
Always research the correct sampling parameters first. This is one of the most important settings, so make sure to find the correct values. For open models, you can find them on Hugging Face in the README and/or in the generation config.
Your parameter selection or settings should leave room for full runs, and you should not restrict things like tokens or number of turns unless specified by the user.
Leave optional settings unset unless the user asks for them. Always confirm the harness, runtime, and sampling parameters before running an evaluation.
Reproducible TOML
You can also use a TOML:
model = "openai/gpt-5-mini"
[env.taskset]
id = "my-task-v1"
split = "test"
[env.agent]
runtime = { type = "subprocess" }
[env.agent.harness]
id = "bash"
[sampling]
temperature = 0.7
uv run vf-eval @ configs/my-eval.toml
Retries
Whole-rollout retry is off by default (max_retries = 0). Each retry starts a fresh rollout. Setting only max_retries retries any captured error up to that cap. Use env.retries for whole-episode retries or env.agent.retries for the agent alone.
Ordered rules override the default budget for matching errors. With a zero default, only explicitly enabled errors retry. For example:
[env.agent.retries]
max_retries = 0
[[env.agent.retries.rules]]
type = "ProviderError"
status_code = [429, "5xx"]
max_retries = 3
[[env.agent.retries.rules]]
type = "SandboxError"
message = 'temporarily unavailable|connection reset'
max_retries = 2
Fields within a rule must all match. type matches the exact recorded exception name, status_code matches any listed status or status class, and message is a regex search (plain text matches a substring). Invalid regexes fail config validation. Omitted match fields match anything. Each rule must explicitly provide max_retries; an omitted budget fails validation.
The first matching rule wins for each error; zero retries excludes it, and exhausted rules never fall through. Unmatched errors share the default max_retries budget. Each retry consumes only the matching rule's budget, or the default budget when no rule matches. Budgets persist across the entire run; the default is not a global cap. An empty rules list uses only the default budget. When an attempt captures multiple errors, the first eligible error triggers the retry; a denied error does not veto other errors. Successful traces' recovered errors do not trigger episode retries.
Output and resume
A run writes to output_dir / run.dir (-o sets output_dir, default outputs; run.dir defaults to the auto-generated run name):
outputs/<env>--<model>--<harness>--<short-id>/
├── configs/eval.json
├── logs/eval.log
└── traces.jsonl
configs/eval.json is the run's resolved config, re-runnable via @. traces.jsonl is one episode per line — the episode's traces plus their shared standing — appended after each episode finishes, so an episode is durable whole or not at all (a torn last line is the whole episode redone on resume).
Resume in place by re-running the run's own saved config with --resume (it re-runs only the missing/errored rollouts; any config drift from the saved run is refused):
uv run vf-eval @ <run-dir>/configs/eval.json --resume
To overwrite a run dir and start fresh instead, use --clean.
Trace inspection
For each representative sample inspect:
taskand prompt fields;branches, assistant messages, tool messages, and stop condition;- named
rewards, aggregatereward, andmetrics; - persisted
infoartifacts; error/errorsand boundary type;- per-call
callsrecords (model, sampling, finish reason, usage, timing, error) linked to the graph; - usage and stage timing;
- token/mask/logprob fields when using the training client.
Classify outcomes:
- Valid completion and correct reward.
- Valid completion with low reward (model/task outcome).
- Truncated completion (budget outcome).
- Captured rollout error (provider, harness, tool, user, runtime, task, or interception).
Do not average these categories together without reporting failure rate.
Metrics interpretation
- Binary rewards support solve rate and pass@k-style analysis.
- Continuous rewards need distributions, quantiles, and per-task/group comparisons.
- Group rewards must be interpreted with their comparison rule and group size.
- Always inspect samples before attributing a delta to model quality.
- Keep taskset, harness, runtime, sampling, and selected task indices fixed across variants.
- Do not overinterpret a tiny smoke run.
When not to use it
- →When the goal is not to reproduce results or evaluate a model/harness combination
Limitations
- →Some harnesses do not offer support to disable tools
- →Do not overinterpret a tiny smoke run
How it compares
This skill provides a structured command-line interface for running and analyzing AI model evaluations with detailed configuration options and output, unlike manual testing or ad-hoc script execution.
Compared to similar skills
evaluate-environments side by side with the closest alternatives in the catalog.
| Skill | Installs | Updated | Safety | Difficulty |
|---|---|---|---|---|
| evaluate-environments (this skill) | 1 | 2mo | Review | Intermediate |
| backtesting-frameworks | 17 | 4mo | No flags | Advanced |
| evaluating-machine-learning-models | 1 | 2mo | Review | Intermediate |
| openjudge | 0 | 6mo | Review | Intermediate |
Try saying
Example prompts that trigger this skill in your AI assistant.
More by PrimeIntellect-ai
View all by PrimeIntellect-ai →You might also like
backtesting-frameworks
wshobson
Build robust backtesting systems for trading strategies with proper handling of look-ahead bias, survivorship bias, and transaction costs. Use when developing trading algorithms, validating strategies, or building backtesting infrastructure.
evaluating-machine-learning-models
jeremylongshore
Build this skill allows AI assistant to evaluate machine learning models using a comprehensive suite of metrics. it should be used when the user requests model performance analysis, validation, or testing. AI assistant can use this skill to assess model accuracy, p... Use when appropriate context detected. Trigger with relevant phrases based on skill purpose.
openjudge
agentscope-ai
>
generate-validation-notebook
monte-carlo-data
Generate SQL validation notebooks for dbt changes. Pass a GitHub PR URL or local dbt repo path.
langsmith-evaluator
dhar174
INVOKE THIS SKILL when building evaluation pipelines for LangSmith. Covers three core components: (1) Creating Evaluators - LLM-as-Judge, custom code; (2) Defining Run Functions - how to capture outputs and trajectories from your agent; (3) Running Evaluations - locally with evaluate() or auto-run v
lofn-qa
LocalSymmetry
Audit, validate, repair, and classify completed Lofn pipeline outputs. Use after music/image/video/story pipeline completion or suspicious partial runs. Do NOT use for creative generation or finalist ranking.