DA

data-engineering-data-pipeline

Provides guidance on designing and implementing robust data ingestion and processing architectures.

Install

mkdir -p .claude/skills/data-engineering-data-pipeline-ngquoctoan2001 && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/13761" && unzip -o skill.zip -d .claude/skills/data-engineering-data-pipeline-ngquoctoan2001 && rm skill.zip

Installs to .claude/skills/data-engineering-data-pipeline-ngquoctoan2001

Activation

This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.

You are a data pipeline architecture expert specializing in scalable, reliable, and cost-effective data pipelines for batch and streaming data processing.
154 charsno explicit “when” trigger
Advanced

Key capabilities

  • Design ETL/ELT, Lambda, Kappa, and Lakehouse architectures
  • Build workflow orchestration with Airflow/Prefect
  • Manage Delta Lake/Iceberg storage with ACID transactions
  • Implement data quality frameworks

How it works

The skill provides expertise in designing various data pipeline architectures, implementing ingestion methods, orchestrating workflows, transforming data, managing storage, and setting up data quality and monitoring frameworks.

Inputs & outputs

You give it
Data pipeline requirements for batch or streaming processing
You get back
Designed data pipeline architecture, implementation code, configuration files, monitoring setup

When to use data-engineering-data-pipeline

  • Design streaming data pipeline
  • Implement dbt transformations
  • Set up data quality checks
  • Optimize pipeline storage costs

About data-engineering-data-pipeline

Offers best practices for ETL/ELT, Lambda, and Kappa architectures. It supports the implementation of orchestration, transformation via dbt/Spark, and data quality monitoring frameworks.

You are a data pipeline architecture expert specializing in scalable, reliable, and cost-effective data pipelines for batch and streaming data processing.

When not to use it

  • When a different domain or tool outside this scope is needed

Limitations

  • The skill is specialized in data pipeline architecture.
  • It focuses on specific tools like Airflow, Prefect, dbt, and Spark.
  • The scope is limited to batch and streaming data processing.

How it compares

This skill offers structured guidance and best practices for building scalable and reliable data pipelines, providing a systematic approach compared to ad-hoc data integration methods.

Compared to similar skills

data-engineering-data-pipeline side by side with the closest alternatives in the catalog.

SkillInstallsUpdatedSafetyDifficulty
data-engineering-data-pipeline (this skill)04moNo flagsAdvanced
sql-queries185moNo flagsIntermediate
senior-data-engineer217moReviewAdvanced
powerbi-modeling116moNo flagsIntermediate

Try saying

Example prompts that trigger this skill in your AI assistant.

You might also like

sql-queries

anthropics

Write correct, performant SQL across all major data warehouse dialects (Snowflake, BigQuery, Databricks, PostgreSQL, etc.). Use when writing queries, optimizing slow SQL, translating between dialects, or building complex analytical queries with CTEs, window functions, or aggregations.

1888

senior-data-engineer

davila7

World-class data engineering skill for building scalable data pipelines, ETL/ELT systems, and data infrastructure. Expertise in Python, SQL, Spark, Airflow, dbt, Kafka, and modern data stack. Includes data modeling, pipeline orchestration, data quality, and DataOps. Use when designing data architectures, building data pipelines, optimizing data workflows, or implementing data governance.

2179

powerbi-modeling

github

Power BI semantic modeling assistant for building optimized data models. Use when working with Power BI semantic models, creating measures, designing star schemas, configuring relationships, implementing RLS, or optimizing model performance. Triggers on queries about DAX calculations, table relationships, dimension/fact table design, naming conventions, model documentation, cardinality, cross-filter direction, calculation groups, and data model best practices. Always connects to the active model first using power-bi-modeling MCP tools to understand the data structure before providing guidance.

1161

spark-optimization

wshobson

Optimize Apache Spark jobs with partitioning, caching, shuffle optimization, and memory tuning. Use when improving Spark performance, debugging slow jobs, or scaling data processing pipelines.

431

data-quality-frameworks

wshobson

Implement data quality validation with Great Expectations, dbt tests, and data contracts. Use when building data quality pipelines, implementing validation rules, or establishing data contracts.

628

fuzzy-matching

dadbodgeoff

Multi-stage fuzzy matching pipeline for entity reconciliation. PostgreSQL trigram pre-filter, salient overlap check, and multi-factor similarity scoring.

825

Search skills

Search the agent skills registry