GE

generating-test-data

Generates type-safe, reproducible test data and fixtures for various test frameworks.

Install

mkdir -p .claude/skills/generating-test-data && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/4770" && unzip -o skill.zip -d .claude/skills/generating-test-data && rm skill.zip

Installs to .claude/skills/generating-test-data

Activation

This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.

Generate realistic test data including edge cases and boundary conditions.
74 charsno explicit “when” trigger
Intermediate

Key capabilities

  • Generate factory functions for data models
  • Create edge case data variants for testing
  • Build relationship factories for connected entity graphs
  • Generate database seed files for integration tests
  • Validate generated data against schemas

How it works

The skill reads project data models to create factory functions that produce valid default instances. It then generates edge case variants and relationship graphs to ensure complete test coverage.

Inputs & outputs

You give it
Data models, TypeScript interfaces, or database schemas
You get back
Factory functions, seed scripts, and fixture files

When to use generating-test-data

  • Generate database seed datasets
  • Create factory functions for unit tests
  • Populate test databases with edge cases
  • Setup fixtures for integration testing

About this skill

Test Data Generator

Overview

Generate realistic, type-safe test data including fixtures, factory functions, seed datasets, and edge case values. Supports Faker.js, Factory Bot patterns, Fishery (TypeScript factories), pytest fixtures, and database seed scripts.

Prerequisites

  • Data generation library installed (Faker.js/@faker-js/faker, Fishery, factory-boy for Python, or JavaFaker)
  • Database schema or TypeScript/Python type definitions for the data models
  • Test framework with fixture support (Jest, pytest, JUnit)
  • Seed management for reproducible random data (faker.seed())
  • Database client for seed data insertion (if generating database fixtures)

Instructions

  1. Read the project's data models, TypeScript interfaces, database schemas, or ORM definitions to understand the shape of all entities.
  2. For each entity, create a factory function that produces a valid default instance:
    • Use Faker methods matched to field semantics (e.g., faker.person.fullName() for names, faker.internet.email() for emails).
    • Provide sensible defaults for required fields.
    • Allow overrides via a partial parameter for test-specific customization.
    • Set a deterministic seed for reproducibility (faker.seed(12345)).
  3. Generate edge case data variants for each entity:
    • Empty values: Empty strings, null, undefined, empty arrays.
    • Boundary values: Maximum string length, integer overflow, zero, negative numbers.
    • Unicode and i18n: Names with accents, CJK characters, RTL text, emoji.
    • Adversarial inputs: SQL injection strings, XSS payloads, excessively long strings.
    • Temporal edge cases: Leap years, timezone boundaries, epoch zero, far-future dates.
  4. Create relationship factories that build connected entity graphs:
    • A user factory that also creates associated addresses and orders.
    • Configurable depth to avoid infinite recursion.
    • Lazy evaluation for optional relationships.
  5. Generate database seed files for integration tests:
    • SQL insert scripts or ORM seed functions.
    • Idempotent operations (use ON CONFLICT or INSERT IF NOT EXISTS).
    • Separate seed sets for different test scenarios (empty state, populated state, edge cases).
  6. Write fixture files in JSON, YAML, or TypeScript for static test data:
    • Group fixtures by test scenario.
    • Include both valid and invalid data sets.
  7. Validate generated data against the schema to ensure factories remain in sync with model changes.

Output

  • Factory function files (one per entity) in test/factories/ or tests/factories/
  • Edge case data collections covering boundaries and adversarial inputs
  • Database seed scripts for integration test environments
  • JSON/YAML fixture files for static test data
  • Factory index file exporting all factories for easy test imports

Error Handling

ErrorCauseSolution
Factory produces invalid dataSchema changed but factory not updatedAdd a validation step that runs the factory output through the schema validator
Duplicate unique valuesFaker generates collisions in small datasetsUse sequential IDs or append a counter; increase Faker's unique retry limit
Database seed fails on foreign keySeed insertion order violates referential integritySort seed operations topologically by dependency; disable FK checks during seeding
Factory recursion overflowCircular relationships (User -> Order -> User)Limit relationship depth; use lazy references; break cycles with ID-only references
Non-deterministic test failuresRandom seed not set consistentlyCall faker.seed() in beforeAll or at factory module level; document seed values

Examples

TypeScript factory with Fishery:

import { Factory } from 'fishery';
import { faker } from '@faker-js/faker';

interface User {
  id: string;
  name: string;
  email: string;
  role: 'admin' | 'user';
  createdAt: Date;
}

export const userFactory = Factory.define<User>(({ sequence }) => ({
  id: `user-${sequence}`,
  name: faker.person.fullName(),
  email: faker.internet.email(),
  role: 'user',
  createdAt: faker.date.past(),
}));

// Usage:
const user = userFactory.build();
const admin = userFactory.build({ role: 'admin' });
const users = userFactory.buildList(10);

pytest fixture factory:

import pytest
from faker import Faker

fake = Faker()
Faker.seed(42)

@pytest.fixture
def make_user():
    def _make_user(**overrides):
        defaults = {
            "name": fake.name(),
            "email": fake.email(),
            "age": fake.random_int(min=18, max=99),
        }
        return {**defaults, **overrides}
    return _make_user

def test_user_validation(make_user):
    user = make_user(age=17)
    assert validate_age(user) is False

Edge case data collection:

export const edgeCases = {
  strings: ['', ' ', '\t\n', 'a'.repeat(10000), '<script>alert(1)</script>',  # 10000: 10 seconds in ms
            "Robert'); DROP TABLE users;--", '\u0000null\u0000byte'],
  numbers: [0, -0, -1, Number.MAX_SAFE_INTEGER, NaN, Infinity, -Infinity],
  dates: [new Date(0), new Date('2024-02-29'), new Date('9999-12-31')],  # 2024: 9999 = configured value
};

Resources

Prerequisites

Data generation libraryDatabase schema or type definitionsTest framework with fixture supportDatabase client for seed data insertion

Limitations

  • Factory recursion overflow on circular relationships
  • Duplicate unique values in small datasets
  • Database seed failures due to referential integrity

How it compares

Unlike manual data creation, this workflow uses deterministic seeds and schema-aware factories to ensure reproducibility and consistency across test environments.

Compared to similar skills

generating-test-data side by side with the closest alternatives in the catalog.

SkillInstallsUpdatedSafetyDifficulty
generating-test-data (this skill)127dReviewIntermediate
lint-and-validate66moReviewBeginner
agent-implementer-sparc-coder16moReviewIntermediate
langsmith-evaluator04moReviewIntermediate

Try saying

Example prompts that trigger this skill in your AI assistant.

More by jeremylongshore

View all by jeremylongshore

analyzing-logs

jeremylongshore

Analyze application logs to detect performance issues, identify error patterns, and improve stability by extracting key insights.

14123

ollama-setup

jeremylongshore

Configure auto-configure Ollama when user needs local LLM deployment, free AI alternatives, or wants to eliminate hosted API costs. Trigger phrases: "install ollama", "local AI", "free LLM", "self-hosted AI", "replace OpenAI", "no API costs". Use when appropriate context detected. Trigger with relevant phrases based on skill purpose.

1167

backtesting-trading-strategies

jeremylongshore

Backtest crypto and traditional trading strategies against historical data. Calculates performance metrics (Sharpe, Sortino, max drawdown), generates equity curves, and optimizes strategy parameters. Use when user wants to test a trading strategy, validate signals, or compare approaches. Trigger with phrases like "backtest strategy", "test trading strategy", "historical performance", "simulate trades", "optimize parameters", or "validate signals".

1071

generating-database-seed-data

jeremylongshore

Process this skill enables AI assistant to generate realistic test data and database seed scripts for development and testing environments. it uses faker libraries to create realistic data, maintains relational integrity, and allows configurable data volumes. u... Use when working with databases or data models. Trigger with phrases like 'database', 'query', or 'schema'.

1033

cursor-codebase-indexing

jeremylongshore

Execute set up and optimize Cursor codebase indexing. Triggers on "cursor index setup", "codebase indexing", "index codebase", "cursor semantic search". Use when working with cursor codebase indexing functionality. Trigger with phrases like "cursor codebase indexing", "cursor indexing", "cursor".

885

testing-mobile-apps

jeremylongshore

Execute mobile app testing on iOS and Android devices/simulators. Use when performing specialized testing. Trigger with phrases like "test mobile app", "run iOS tests", or "validate Android functionality".

810

You might also like

lint-and-validate

davila7

Automatic quality control, linting, and static analysis procedures. Use after every code modification to ensure syntax correctness and project standards. Triggers onKeywords: lint, format, check, validate, types, static analysis.

68

agent-implementer-sparc-coder

ruvnet

Agent skill for implementer-sparc-coder - invoke with $agent-implementer-sparc-coder

11

langsmith-evaluator

dhar174

INVOKE THIS SKILL when building evaluation pipelines for LangSmith. Covers three core components: (1) Creating Evaluators - LLM-as-Judge, custom code; (2) Defining Run Functions - how to capture outputs and trajectories from your agent; (3) Running Evaluations - locally with evaluate() or auto-run v

00

assets

salazarsebas

Stellar Assets (classic) + trustlines + Stellar Asset Contract (SAC) bridge to Soroban. Covers asset issuance, distribution, authorization flags, clawback, regulated assets, trustline management, and the SAC interop layer that exposes classic assets as Soroban tokens. Use when tokenizing real-world

00

data

salazarsebas

Querying Stellar chain data via Stellar RPC (preferred) and Horizon (legacy). Covers RPC JSON-RPC methods, Horizon REST endpoints, streaming, pagination, historical queries, Hubble/Galexie for deep history, and the RPC/Horizon migration story. Use when reading balances, transactions, operations, led

00

mcp-builder

anthropics

Guide for creating high-quality MCP (Model Context Protocol) servers that enable LLMs to interact with external services through well-designed tools. Use when building MCP servers to integrate external APIs or services, whether in Python (FastMCP) or Node/TypeScript (MCP SDK).

136215

Search skills

Search the agent skills registry