apache-spark
Enables large-scale data processing and analytics using Apache Spark frameworks.
Install
mkdir -p .claude/skills/apache-spark && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/16761" && unzip -o skill.zip -d .claude/skills/apache-spark && rm skill.zipInstalls to .claude/skills/apache-spark
Activation
This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.
Processes large-scale data with Apache Spark using DataFrames, RDDs, and Spark SQL. Use for big data ETL and analytics.Key capabilities
- →process terabyte-scale data
- →perform ETL operations
- →conduct interactive analytics
- →engineer features for machine learning
- →execute DataFrame operations
- →run Spark SQL queries
How it works
The skill initializes a SparkSession, reads CSV data into a DataFrame, and then performs operations like filtering, adding columns, and aggregating data.
Inputs & outputs
When to use apache-spark
- →ETL pipelines
- →Large-scale data processing
- →Interactive analytics
- →ML feature engineering
About this skill
Apache Spark
Large-scale distributed data processing with Spark.
Quick Start
pip install pyspark
pyspark # interactive shell
When to Use
- Processing datasets too big for a single machine
- Distributed ETL and analytics
- Streaming (Structured Streaming)
- Machine learning at scale (MLlib)
Best Practices
DataFrames & SQL
- Prefer DataFrames over RDDs for most work
- Use Spark SQL for familiar declarative queries
- Cache intermediate DataFrames when reused
- Use
broadcastfor small join tables
Optimization
- Avoid shuffles; use partitioning and bucketing
- Filter and select early to reduce data
- Use
repartition/coalescedeliberately - Monitor stages and shuffle in the UI
Joins & Aggregations
- Broadcast small tables with
broadcast() - Join on bucketed/partitioned keys
- Use
groupBywith aggregation functions - Prefer window functions carefully
Resource Tuning
- Set
spark.executor.memoryand cores - Tune
spark.sql.shuffle.partitions - Use dynamic allocation
- Balance partitions to avoid skew
Dependencies
pip install pyspark
Examples
from pyspark.sql import SparkSession
from pyspark.sql import functions as F
spark = SparkSession.builder.appName("etl").getOrCreate()
df = spark.read.parquet("s3://bucket/events")
print(df.printSchema())
# ETL with filters and aggregations
result = (
df.filter(F.col("status") == "paid")
.groupBy("region")
.agg(F.sum("amount").alias("total"), F.count("*").alias("orders"))
.orderBy(F.desc("total"))
)
result.show()
# Broadcast join
from pyspark.sql.functions import broadcast
users = spark.read.parquet("users.parquet").cache()
joined = df.join(broadcast(users), "user_id")
-- Spark SQL
SELECT region, sum(amount) AS total
FROM events
WHERE status = 'paid'
GROUP BY region
ORDER BY total DESC
Step-by-Step
- Start a SparkSession with tuned config.
- Load data (parquet/json/table).
- Transform with DataFrames/SQL.
- Optimize: broadcast, partitioning, caching.
- Write results in a columnar format.
- Monitor the Spark UI for skew and shuffles.
- Tune executors and partitions.
- Schedule jobs (Airflow/spark-submit).
Validation
- Results match a small reference computation
- Shuffles are minimized
- No task skew (balanced partitions)
- Caching speeds up reused data
- Job completes within the budget
Troubleshooting
- OOM: reduce executor memory or partition data.
- Skew: salt keys or repartition.
- Slow shuffles: tune shuffle.partitions and bucketing.
When not to use it
- →when data processing is not terabyte-scale
- →when ETL pipelines are not required
- →when interactive analytics are not needed
Limitations
- →limited to large-scale data processing
- →focused on ETL pipelines
- →designed for interactive analytics
How it compares
This skill provides a unified analytics engine for large-scale data processing, unlike manual data handling that would require separate tools for different data sizes and operations.
Compared to similar skills
apache-spark side by side with the closest alternatives in the catalog.
| Skill | Installs | Updated | Safety | Difficulty |
|---|---|---|---|---|
| apache-spark (this skill) | 0 | 3mo | No flags | Intermediate |
| hugging-face-datasets | 1 | 8mo | Review | Intermediate |
| quant-analyst | 103 | 4mo | No flags | Advanced |
| umap-learn | 6 | 3mo | Review | Intermediate |
Try saying
Example prompts that trigger this skill in your AI assistant.
More by ssrjkk
View all by ssrjkk →You might also like
hugging-face-datasets
patchy631
Create and manage datasets on Hugging Face Hub. Supports initializing repos, defining configs/system prompts, streaming row updates, and SQL-based dataset querying/transformation. Designed to work alongside HF MCP server for comprehensive dataset workflows.
quant-analyst
zenobi-us
Expert quantitative analyst specializing in financial modeling, algorithmic trading, and risk analytics. Masters statistical methods, derivatives pricing, and high-frequency trading with focus on mathematical rigor, performance optimization, and profitable strategy development.
umap-learn
K-Dense-AI
UMAP dimensionality reduction. Fast nonlinear manifold learning for 2D/3D visualization, clustering preprocessing (HDBSCAN), supervised/parametric UMAP, for high-dimensional data.
senior-data-engineer
davila7
World-class data engineering skill for building scalable data pipelines, ETL/ELT systems, and data infrastructure. Expertise in Python, SQL, Spark, Airflow, dbt, Kafka, and modern data stack. Includes data modeling, pipeline orchestration, data quality, and DataOps. Use when designing data architectures, building data pipelines, optimizing data workflows, or implementing data governance.
embedding-strategies
wshobson
Select and optimize embedding models for semantic search and RAG applications. Use when choosing embedding models, implementing chunking strategies, or optimizing embedding quality for specific domains.
building-automl-pipelines
jeremylongshore
Build automated machine learning pipelines, including feature engineering, model selection, and performance evaluation.