synthesize-data
Creates synthetic data pairs based on task specifications to power benchmarks.
Install
mkdir -p .claude/skills/synthesize-data && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/14179" && unzip -o skill.zip -d .claude/skills/synthesize-data && rm skill.zipInstalls to .claude/skills/synthesize-data
Activation
This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.
Generate synthetic benchmark examples from a task spec. Invoked from setup-data when there is no data, or directly when the user asks to generate examples. Produces a JSONL file in .data/<benchmark-name>/.Key capabilities
- →Read benchmark specifications
- →Generate synthetic text and expected output pairs
- →Validate generated results against field schema
- →Save generated data to a JSONL file
- →Update benchmark spec to point to generated data
- →Warn about small datasets for statistical reliability
How it works
The skill reads a benchmark specification, builds a generation prompt, calls a generator model, validates results, and saves them to a JSONL file.
Inputs & outputs
When to use synthesize-data
- →Generate test data for benchmarks
- →Create synthetic evaluation examples
- →Validate benchmark task specs
About synthesize-data
Reads benchmark requirements and generates sample input-output pairs. Validates output against the specification, ensuring robust testing data.
Generate synthetic benchmark examples from a task spec. Invoked from setup-data when there is no data, or directly when the user asks to generate examples. Produces a JSONL file in .data/<benchmark-name>/.
When not to use it
- →When `task.input.type` is `image`, `document`, or `pdf`
- →When real images are required for the benchmark
Limitations
- →The skill cannot generate images, documents, or PDFs.
- →The skill relies on an internal generator model that is not exposed to the user.
How it compares
This skill automates the creation of synthetic benchmark data based on a specification, providing structured and validated examples, which differs from manually crafting test data.
Compared to similar skills
synthesize-data side by side with the closest alternatives in the catalog.
| Skill | Installs | Updated | Safety | Difficulty |
|---|---|---|---|---|
| synthesize-data (this skill) | 0 | 3mo | Review | Intermediate |
| data-engineering | 13 | 9mo | Review | Advanced |
| crawl4ai | 21 | 10mo | Review | Intermediate |
| data-cleaning-pipeline | 13 | 7mo | Review | Intermediate |
Try saying
Example prompts that trigger this skill in your AI assistant.
More by surus-lat
View all by surus-lat →You might also like
data-engineering
pluginagentmarketplace
ETL pipelines, Apache Spark, data warehousing, and big data processing. Use for building data pipelines, processing large datasets, or data infrastructure.
crawl4ai
basher83
This skill should be used when users need to scrape websites, extract structured data, handle JavaScript-heavy pages, crawl multiple URLs, or build automated web data pipelines. Includes optimized extraction patterns with schema generation for efficient, LLM-free extraction.
data-cleaning-pipeline
aj-geddes
Build robust processes for data cleaning, missing value imputation, outlier handling, and data transformation for data preprocessing, data quality, and data pipeline automation
paddle-ocr-validation
jgtolentino
PaddleOCR-based receipt and BIR form extraction with validation
apify
vm0-ai
Web scraping and automation platform with pre-built Actors for common tasks
ocr
trpc-group
Extract text from images using Tesseract OCR