Enables large-scale data processing and analytics using Apache Spark frameworks.

Install

mkdir -p .claude/skills/apache-spark && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/16761" && unzip -o skill.zip -d .claude/skills/apache-spark && rm skill.zip

Installs to .claude/skills/apache-spark

Activation

This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.

Processes large-scale data with Apache Spark using DataFrames, RDDs, and Spark SQL. Use for big data ETL and analytics.
119 chars✓ has a “when” trigger
Intermediate

Key capabilities

  • →process terabyte-scale data
  • →perform ETL operations
  • →conduct interactive analytics
  • →engineer features for machine learning
  • →execute DataFrame operations
  • →run Spark SQL queries

How it works

The skill initializes a SparkSession, reads CSV data into a DataFrame, and then performs operations like filtering, adding columns, and aggregating data.

Inputs & outputs

You give it
CSV files from 'data/*.csv'
You get back
aggregated sum of 'amount' grouped by 'category'

When to use apache-spark

  • →ETL pipelines
  • →Large-scale data processing
  • →Interactive analytics
  • →ML feature engineering

About this skill

Apache Spark

Large-scale distributed data processing with Spark.

Quick Start

pip install pyspark
pyspark  # interactive shell

When to Use

  • Processing datasets too big for a single machine
  • Distributed ETL and analytics
  • Streaming (Structured Streaming)
  • Machine learning at scale (MLlib)

Best Practices

DataFrames & SQL

  • Prefer DataFrames over RDDs for most work
  • Use Spark SQL for familiar declarative queries
  • Cache intermediate DataFrames when reused
  • Use broadcast for small join tables

Optimization

  • Avoid shuffles; use partitioning and bucketing
  • Filter and select early to reduce data
  • Use repartition/coalesce deliberately
  • Monitor stages and shuffle in the UI

Joins & Aggregations

  • Broadcast small tables with broadcast()
  • Join on bucketed/partitioned keys
  • Use groupBy with aggregation functions
  • Prefer window functions carefully

Resource Tuning

  • Set spark.executor.memory and cores
  • Tune spark.sql.shuffle.partitions
  • Use dynamic allocation
  • Balance partitions to avoid skew

Dependencies

pip install pyspark

Examples

from pyspark.sql import SparkSession
from pyspark.sql import functions as F

spark = SparkSession.builder.appName("etl").getOrCreate()

df = spark.read.parquet("s3://bucket/events")
print(df.printSchema())
# ETL with filters and aggregations
result = (
    df.filter(F.col("status") == "paid")
      .groupBy("region")
      .agg(F.sum("amount").alias("total"), F.count("*").alias("orders"))
      .orderBy(F.desc("total"))
)
result.show()
# Broadcast join
from pyspark.sql.functions import broadcast

users = spark.read.parquet("users.parquet").cache()
joined = df.join(broadcast(users), "user_id")
-- Spark SQL
SELECT region, sum(amount) AS total
FROM events
WHERE status = 'paid'
GROUP BY region
ORDER BY total DESC

Step-by-Step

  1. Start a SparkSession with tuned config.
  2. Load data (parquet/json/table).
  3. Transform with DataFrames/SQL.
  4. Optimize: broadcast, partitioning, caching.
  5. Write results in a columnar format.
  6. Monitor the Spark UI for skew and shuffles.
  7. Tune executors and partitions.
  8. Schedule jobs (Airflow/spark-submit).

Validation

  1. Results match a small reference computation
  2. Shuffles are minimized
  3. No task skew (balanced partitions)
  4. Caching speeds up reused data
  5. Job completes within the budget

Troubleshooting

  • OOM: reduce executor memory or partition data.
  • Skew: salt keys or repartition.
  • Slow shuffles: tune shuffle.partitions and bucketing.

When not to use it

  • →when data processing is not terabyte-scale
  • →when ETL pipelines are not required
  • →when interactive analytics are not needed

Limitations

  • →limited to large-scale data processing
  • →focused on ETL pipelines
  • →designed for interactive analytics

How it compares

This skill provides a unified analytics engine for large-scale data processing, unlike manual data handling that would require separate tools for different data sizes and operations.

Compared to similar skills

apache-spark side by side with the closest alternatives in the catalog.

SkillInstallsUpdatedSafetyDifficulty
apache-spark (this skill)03moNo flagsIntermediate
hugging-face-datasets18moReviewIntermediate
quant-analyst1034moNo flagsAdvanced
umap-learn63moReviewIntermediate

Try saying

Example prompts that trigger this skill in your AI assistant.

You might also like

hugging-face-datasets

patchy631

Create and manage datasets on Hugging Face Hub. Supports initializing repos, defining configs/system prompts, streaming row updates, and SQL-based dataset querying/transformation. Designed to work alongside HF MCP server for comprehensive dataset workflows.

14

quant-analyst

zenobi-us

Expert quantitative analyst specializing in financial modeling, algorithmic trading, and risk analytics. Masters statistical methods, derivatives pricing, and high-frequency trading with focus on mathematical rigor, performance optimization, and profitable strategy development.

103355

umap-learn

K-Dense-AI

UMAP dimensionality reduction. Fast nonlinear manifold learning for 2D/3D visualization, clustering preprocessing (HDBSCAN), supervised/parametric UMAP, for high-dimensional data.

6100

senior-data-engineer

davila7

World-class data engineering skill for building scalable data pipelines, ETL/ELT systems, and data infrastructure. Expertise in Python, SQL, Spark, Airflow, dbt, Kafka, and modern data stack. Includes data modeling, pipeline orchestration, data quality, and DataOps. Use when designing data architectures, building data pipelines, optimizing data workflows, or implementing data governance.

2179

embedding-strategies

wshobson

Select and optimize embedding models for semantic search and RAG applications. Use when choosing embedding models, implementing chunking strategies, or optimizing embedding quality for specific domains.

890

building-automl-pipelines

jeremylongshore

Build automated machine learning pipelines, including feature engineering, model selection, and performance evaluation.

688

Search skills

Search the agent skills registry