dataeng-codebase-analyst
Inspects data pipelines, jobs, and monitoring assets to identify technical debt and structural gaps.
Install
mkdir -p .claude/skills/dataeng-codebase-analyst && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/11616" && unzip -o skill.zip -d .claude/skills/dataeng-codebase-analyst && rm skill.zipInstalls to .claude/skills/dataeng-codebase-analyst
Activation
This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.
Analyze existing Data Engineering codebase — data pipeline, reproducibility, artifacts, and operational flowKey capabilities
- →Identify data pipeline patterns
- →Assess pipeline reproducibility
- →Evaluate artifact management
- →Pinpoint testing gaps
- →Uncover technical debt
How it works
The skill inspects files like main.py, databricks.yml, params.yaml, runtime artifact expectations, notebooks, and monitoring assets.
Inputs & outputs
When to use dataeng-codebase-analyst
- →Analyze data pipeline patterns
- →Identify technical debt in pipeline
- →Review operational flow for reproduction
About this skill
Skill: Codebase Analyst
Purpose
Analyze an existing Data Engineering codebase to identify data pipeline patterns, pipeline reproducibility, artifact management, testing gaps, and technical debt.
What to inspect
main.pydatabricks.yml,Databricks job run metadata,params.yamldatabricks.yml- runtime artifact expectations
- notebooks and extracted analytical helpers
- monitoring and operational assets
When not to use it
- →When the codebase is not related to Data Engineering
- →When the goal is not to analyze data pipeline, reproducibility, artifacts, or operational flow
Limitations
- →Limited to analyzing Data Engineering codebases
- →Focuses on specific files and aspects like main.py, databricks.yml, and runtime artifacts
How it compares
This skill automates the inspection of specific Data Engineering codebase components, unlike a manual review.
Compared to similar skills
dataeng-codebase-analyst side by side with the closest alternatives in the catalog.
| Skill | Installs | Updated | Safety | Difficulty |
|---|---|---|---|---|
| dataeng-codebase-analyst (this skill) | 0 | 3mo | No flags | Intermediate |
| spark-optimization | 4 | 2mo | No flags | Advanced |
| data-quality-monitor-designer | 0 | 1mo | No flags | Advanced |
| feast-user-guide | 0 | 1mo | Review | Intermediate |
Try saying
Example prompts that trigger this skill in your AI assistant.
You might also like
spark-optimization
wshobson
Optimize Apache Spark jobs with partitioning, caching, shuffle optimization, and memory tuning. Use when improving Spark performance, debugging slow jobs, or scaling data processing pipelines.
data-quality-monitor-designer
nguyenpv1980-wq
Design data-quality monitoring for production pipelines and stores — checks across the six dimensions (freshness, completeness/volume, uniqueness, validity, consistency/referential integrity, distribution drift), placed at the right pipeline stage (ingest, transform, serving), each with severity, an
feast-user-guide
feast-dev
Guide for working with Feast (Feature Store) — defining features, configuring feature_store.yaml, retrieving features online/offline, using the CLI, and building RAG retrieval pipelines. Use when the user asks about creating entities, feature views, on-demand feature views, stream feature views, fea
gcp-spark
ironkid90
|
senior-data-engineer
davila7
World-class data engineering skill for building scalable data pipelines, ETL/ELT systems, and data infrastructure. Expertise in Python, SQL, Spark, Airflow, dbt, Kafka, and modern data stack. Includes data modeling, pipeline orchestration, data quality, and DataOps. Use when designing data architectures, building data pipelines, optimizing data workflows, or implementing data governance.
senior-data-scientist
davila7
World-class data science skill for statistical modeling, experimentation, causal inference, and advanced analytics. Expertise in Python (NumPy, Pandas, Scikit-learn), R, SQL, statistical methods, A/B testing, time series, and business intelligence. Includes experiment design, feature engineering, model evaluation, and stakeholder communication. Use when designing experiments, building predictive models, performing causal analysis, or driving data-driven decisions.