DA

dataeng-codebase-analyst

Inspects data pipelines, jobs, and monitoring assets to identify technical debt and structural gaps.

Install

mkdir -p .claude/skills/dataeng-codebase-analyst && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/11616" && unzip -o skill.zip -d .claude/skills/dataeng-codebase-analyst && rm skill.zip

Installs to .claude/skills/dataeng-codebase-analyst

Activation

This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.

Analyze existing Data Engineering codebase — data pipeline, reproducibility, artifacts, and operational flow
108 charsno explicit “when” trigger
Intermediate

Key capabilities

  • Identify data pipeline patterns
  • Assess pipeline reproducibility
  • Evaluate artifact management
  • Pinpoint testing gaps
  • Uncover technical debt

How it works

The skill inspects files like main.py, databricks.yml, params.yaml, runtime artifact expectations, notebooks, and monitoring assets.

Inputs & outputs

You give it
An existing Data Engineering codebase
You get back
Identified data pipeline patterns, reproducibility issues, artifact management status, testing gaps, and technical debt

When to use dataeng-codebase-analyst

  • Analyze data pipeline patterns
  • Identify technical debt in pipeline
  • Review operational flow for reproduction

About this skill

Skill: Codebase Analyst

Purpose

Analyze an existing Data Engineering codebase to identify data pipeline patterns, pipeline reproducibility, artifact management, testing gaps, and technical debt.

What to inspect

  • main.py
  • databricks.yml, Databricks job run metadata, params.yaml
  • databricks.yml
  • runtime artifact expectations
  • notebooks and extracted analytical helpers
  • monitoring and operational assets

When not to use it

  • When the codebase is not related to Data Engineering
  • When the goal is not to analyze data pipeline, reproducibility, artifacts, or operational flow

Limitations

  • Limited to analyzing Data Engineering codebases
  • Focuses on specific files and aspects like main.py, databricks.yml, and runtime artifacts

How it compares

This skill automates the inspection of specific Data Engineering codebase components, unlike a manual review.

Compared to similar skills

dataeng-codebase-analyst side by side with the closest alternatives in the catalog.

SkillInstallsUpdatedSafetyDifficulty
dataeng-codebase-analyst (this skill)03moNo flagsIntermediate
spark-optimization42moNo flagsAdvanced
data-quality-monitor-designer01moNo flagsAdvanced
feast-user-guide01moReviewIntermediate

Try saying

Example prompts that trigger this skill in your AI assistant.

You might also like

spark-optimization

wshobson

Optimize Apache Spark jobs with partitioning, caching, shuffle optimization, and memory tuning. Use when improving Spark performance, debugging slow jobs, or scaling data processing pipelines.

431

data-quality-monitor-designer

nguyenpv1980-wq

Design data-quality monitoring for production pipelines and stores — checks across the six dimensions (freshness, completeness/volume, uniqueness, validity, consistency/referential integrity, distribution drift), placed at the right pipeline stage (ingest, transform, serving), each with severity, an

00

feast-user-guide

feast-dev

Guide for working with Feast (Feature Store) — defining features, configuring feature_store.yaml, retrieving features online/offline, using the CLI, and building RAG retrieval pipelines. Use when the user asks about creating entities, feature views, on-demand feature views, stream feature views, fea

00

gcp-spark

ironkid90

|

00

senior-data-engineer

davila7

World-class data engineering skill for building scalable data pipelines, ETL/ELT systems, and data infrastructure. Expertise in Python, SQL, Spark, Airflow, dbt, Kafka, and modern data stack. Includes data modeling, pipeline orchestration, data quality, and DataOps. Use when designing data architectures, building data pipelines, optimizing data workflows, or implementing data governance.

2179

senior-data-scientist

davila7

World-class data science skill for statistical modeling, experimentation, causal inference, and advanced analytics. Expertise in Python (NumPy, Pandas, Scikit-learn), R, SQL, statistical methods, A/B testing, time series, and business intelligence. Includes experiment design, feature engineering, model evaluation, and stakeholder communication. Use when designing experiments, building predictive models, performing causal analysis, or driving data-driven decisions.

952

Search skills

Search the agent skills registry