data-engineering-data-pipeline
Provides guidance on designing and implementing robust data ingestion and processing architectures.
Install
mkdir -p .claude/skills/data-engineering-data-pipeline-ngquoctoan2001 && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/13761" && unzip -o skill.zip -d .claude/skills/data-engineering-data-pipeline-ngquoctoan2001 && rm skill.zipInstalls to .claude/skills/data-engineering-data-pipeline-ngquoctoan2001
Activation
This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.
You are a data pipeline architecture expert specializing in scalable, reliable, and cost-effective data pipelines for batch and streaming data processing.Key capabilities
- →Design ETL/ELT, Lambda, Kappa, and Lakehouse architectures
- →Build workflow orchestration with Airflow/Prefect
- →Manage Delta Lake/Iceberg storage with ACID transactions
- →Implement data quality frameworks
How it works
The skill provides expertise in designing various data pipeline architectures, implementing ingestion methods, orchestrating workflows, transforming data, managing storage, and setting up data quality and monitoring frameworks.
Inputs & outputs
When to use data-engineering-data-pipeline
- →Design streaming data pipeline
- →Implement dbt transformations
- →Set up data quality checks
- →Optimize pipeline storage costs
About data-engineering-data-pipeline
Offers best practices for ETL/ELT, Lambda, and Kappa architectures. It supports the implementation of orchestration, transformation via dbt/Spark, and data quality monitoring frameworks.
You are a data pipeline architecture expert specializing in scalable, reliable, and cost-effective data pipelines for batch and streaming data processing.
When not to use it
- →When a different domain or tool outside this scope is needed
Limitations
- →The skill is specialized in data pipeline architecture.
- →It focuses on specific tools like Airflow, Prefect, dbt, and Spark.
- →The scope is limited to batch and streaming data processing.
How it compares
This skill offers structured guidance and best practices for building scalable and reliable data pipelines, providing a systematic approach compared to ad-hoc data integration methods.
Compared to similar skills
data-engineering-data-pipeline side by side with the closest alternatives in the catalog.
| Skill | Installs | Updated | Safety | Difficulty |
|---|---|---|---|---|
| data-engineering-data-pipeline (this skill) | 0 | 4mo | No flags | Advanced |
| sql-queries | 18 | 5mo | No flags | Intermediate |
| senior-data-engineer | 21 | 7mo | Review | Advanced |
| powerbi-modeling | 11 | 6mo | No flags | Intermediate |
Try saying
Example prompts that trigger this skill in your AI assistant.
You might also like
sql-queries
anthropics
Write correct, performant SQL across all major data warehouse dialects (Snowflake, BigQuery, Databricks, PostgreSQL, etc.). Use when writing queries, optimizing slow SQL, translating between dialects, or building complex analytical queries with CTEs, window functions, or aggregations.
senior-data-engineer
davila7
World-class data engineering skill for building scalable data pipelines, ETL/ELT systems, and data infrastructure. Expertise in Python, SQL, Spark, Airflow, dbt, Kafka, and modern data stack. Includes data modeling, pipeline orchestration, data quality, and DataOps. Use when designing data architectures, building data pipelines, optimizing data workflows, or implementing data governance.
powerbi-modeling
github
Power BI semantic modeling assistant for building optimized data models. Use when working with Power BI semantic models, creating measures, designing star schemas, configuring relationships, implementing RLS, or optimizing model performance. Triggers on queries about DAX calculations, table relationships, dimension/fact table design, naming conventions, model documentation, cardinality, cross-filter direction, calculation groups, and data model best practices. Always connects to the active model first using power-bi-modeling MCP tools to understand the data structure before providing guidance.
spark-optimization
wshobson
Optimize Apache Spark jobs with partitioning, caching, shuffle optimization, and memory tuning. Use when improving Spark performance, debugging slow jobs, or scaling data processing pipelines.
data-quality-frameworks
wshobson
Implement data quality validation with Great Expectations, dbt tests, and data contracts. Use when building data quality pipelines, implementing validation rules, or establishing data contracts.
fuzzy-matching
dadbodgeoff
Multi-stage fuzzy matching pipeline for entity reconciliation. PostgreSQL trigram pre-filter, salient overlap check, and multi-factor similarity scoring.