Integrates Terminal-Bench and Harbor for benchmarking AI agent performance and failure analysis.

Install

mkdir -p .claude/skills/tbench && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/8589" && unzip -o skill.zip -d .claude/skills/tbench && rm skill.zip

Installs to .claude/skills/tbench

Activation

This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.

Terminal-Bench integration for Mux agent benchmarking and failure analysis
74 charsno explicit “when” trigger
Intermediate

Key capabilities

  • →Run full benchmark suites for agent evaluation
  • →Execute specific benchmark tasks
  • →Run benchmarks with specific models and thinking levels
  • →Utilize Daytona cloud sandboxes for parallel execution
  • →Analyze failure rates of agents
  • →Configure global timeouts for benchmark runs

How it works

It integrates with Terminal-Bench 2.0 and Harbor to execute benchmark suites for agent evaluation, allowing configuration of tasks, models, and execution environments. It also provides tools for analyzing the results.

Inputs & outputs

You give it
Benchmark commands with optional environment variables and arguments
You get back
Benchmark results, including success/failure rates and performance metrics

When to use tbench

  • →Run benchmark suite
  • →Analyze agent performance
  • →Use cloud sandboxes for faster testing

About tbench

Executes benchmark suites for agent evaluation using local Docker or cloud sandboxes. It provides tools for running specific tasks and analyzing performance results.

Terminal-Bench integration for Mux agent benchmarking and failure analysis

When not to use it

  • →When the task requires a per-task timeout configuration instead of a global one

Limitations

  • →The default global timeout is 30 minutes, which may not be sufficient for all tasks.
  • →The skill prefers global timeout defaults over per-task configuration.
  • →The analysis script requires `bq CLI` for Mux results and `git` for leaderboard data.

How it compares

This skill automates and standardizes agent benchmarking using specific tools and configurations, offering a structured approach compared to ad-hoc testing.

Compared to similar skills

tbench side by side with the closest alternatives in the catalog.

SkillInstallsUpdatedSafetyDifficulty
tbench (this skill)04moReviewIntermediate
chaos-scenario03moReviewAdvanced
omnidocbench-eval-helper05moReviewAdvanced
mlops-engineer35moNo flagsAdvanced

Try saying

Example prompts that trigger this skill in your AI assistant.

You might also like

chaos-scenario

petercort

Use when authoring, running, or reviewing chaos engineering experiments in this monorepo. Covers steady-state hypothesis, fault injection (service kill, latency, network partition, overload), result recording to CSV, and cleanup/restore. Triggers: "add chaos scenario", "new chaos test", "inject faul

00

omnidocbench-eval-helper

opendatalab

Help users deploy, validate, run, and parse OmniDocBench evaluations. Use this skill whenever the user mentions OmniDocBench, document parsing/OCR benchmark scoring, MinerU or other model evaluation on OmniDocBench, CDM formula metrics, end2end/md2md configs, Docker/conda deployment, remote SSH/H-cl

00

mlops-engineer

sickn33

Build comprehensive ML pipelines, experiment tracking, and model registries with MLflow, Kubeflow, and modern MLOps tools. Implements automated training, deployment, and monitoring across cloud platforms. Use PROACTIVELY for ML infrastructure, experiment management, or pipeline automation.

333

senior-devops

davila7

Comprehensive DevOps skill for CI/CD, infrastructure automation, containerization, and cloud platforms (AWS, GCP, Azure). Includes pipeline setup, infrastructure as code, deployment automation, and monitoring. Use when setting up pipelines, deploying applications, managing infrastructure, implementing monitoring, or optimizing deployment processes.

720

log-analyzer

mikopbx

Анализ логов Docker контейнера для диагностики проблем и мониторинга здоровья системы. Использовать при отладке ошибок, отслеживании процессов воркеров, исследовании проблем API или мониторинге поведения системы после тестов.

213

server-management

davila7

Server management principles and decision-making. Process management, monitoring strategy, and scaling decisions. Teaches thinking, not commands.

113

Search skills

Search the agent skills registry