Integrates Terminal-Bench and Harbor for benchmarking AI agent performance and failure analysis.
Install
mkdir -p .claude/skills/tbench && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/8589" && unzip -o skill.zip -d .claude/skills/tbench && rm skill.zipInstalls to .claude/skills/tbench
Activation
This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.
Terminal-Bench integration for Mux agent benchmarking and failure analysisKey capabilities
- →Run full benchmark suites for agent evaluation
- →Execute specific benchmark tasks
- →Run benchmarks with specific models and thinking levels
- →Utilize Daytona cloud sandboxes for parallel execution
- →Analyze failure rates of agents
- →Configure global timeouts for benchmark runs
How it works
It integrates with Terminal-Bench 2.0 and Harbor to execute benchmark suites for agent evaluation, allowing configuration of tasks, models, and execution environments. It also provides tools for analyzing the results.
Inputs & outputs
When to use tbench
- →Run benchmark suite
- →Analyze agent performance
- →Use cloud sandboxes for faster testing
About tbench
Executes benchmark suites for agent evaluation using local Docker or cloud sandboxes. It provides tools for running specific tasks and analyzing performance results.
Terminal-Bench integration for Mux agent benchmarking and failure analysis
When not to use it
- →When the task requires a per-task timeout configuration instead of a global one
Limitations
- →The default global timeout is 30 minutes, which may not be sufficient for all tasks.
- →The skill prefers global timeout defaults over per-task configuration.
- →The analysis script requires `bq CLI` for Mux results and `git` for leaderboard data.
How it compares
This skill automates and standardizes agent benchmarking using specific tools and configurations, offering a structured approach compared to ad-hoc testing.
Compared to similar skills
tbench side by side with the closest alternatives in the catalog.
| Skill | Installs | Updated | Safety | Difficulty |
|---|---|---|---|---|
| tbench (this skill) | 0 | 4mo | Review | Intermediate |
| chaos-scenario | 0 | 3mo | Review | Advanced |
| omnidocbench-eval-helper | 0 | 5mo | Review | Advanced |
| mlops-engineer | 3 | 5mo | No flags | Advanced |
Try saying
Example prompts that trigger this skill in your AI assistant.
More by coder
View all by coder →You might also like
chaos-scenario
petercort
Use when authoring, running, or reviewing chaos engineering experiments in this monorepo. Covers steady-state hypothesis, fault injection (service kill, latency, network partition, overload), result recording to CSV, and cleanup/restore. Triggers: "add chaos scenario", "new chaos test", "inject faul
omnidocbench-eval-helper
opendatalab
Help users deploy, validate, run, and parse OmniDocBench evaluations. Use this skill whenever the user mentions OmniDocBench, document parsing/OCR benchmark scoring, MinerU or other model evaluation on OmniDocBench, CDM formula metrics, end2end/md2md configs, Docker/conda deployment, remote SSH/H-cl
mlops-engineer
sickn33
Build comprehensive ML pipelines, experiment tracking, and model registries with MLflow, Kubeflow, and modern MLOps tools. Implements automated training, deployment, and monitoring across cloud platforms. Use PROACTIVELY for ML infrastructure, experiment management, or pipeline automation.
senior-devops
davila7
Comprehensive DevOps skill for CI/CD, infrastructure automation, containerization, and cloud platforms (AWS, GCP, Azure). Includes pipeline setup, infrastructure as code, deployment automation, and monitoring. Use when setting up pipelines, deploying applications, managing infrastructure, implementing monitoring, or optimizing deployment processes.
log-analyzer
mikopbx
Анализ логов Docker контейнера для диагностики проблем и мониторинга здоровья системы. Использовать при отладке ошибок, отслеживании процессов воркеров, исследовании проблем API или мониторинге поведения системы после тестов.
server-management
davila7
Server management principles and decision-making. Process management, monitoring strategy, and scaling decisions. Teaches thinking, not commands.