BI

bio-chip-seq-super-enhancers

Detects super-enhancers from H3K27ac/MED1/BRD4 ChIP-seq data to study cell-identity gene regulation.

Install

mkdir -p .claude/skills/bio-chip-seq-super-enhancers && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/5814" && unzip -o skill.zip -d .claude/skills/bio-chip-seq-super-enhancers && rm skill.zip

Installs to .claude/skills/bio-chip-seq-super-enhancers

Activation

This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.

Identifies super-enhancers from H3K27ac, MED1, or BRD4 ChIP-seq using ROSE, ROSE2, LILY, HOMER -style super, and ENCODE dELS cross-referencing. Handles peak stitching parameters, ranking choices, hockey-stick inflection, marker choice (H3K27ac vs MED1/BRD4), and cross-condition comparison with spike-in normalization. Constructs core regulatory circuitry (Saint-Andre 2016) from SE-encoded TFs. Use when identifying cell-identity / cancer-associated regulatory domains, comparing super-enhancers between conditions, identifying master transcription factor networks, or predicting BET-inhibitor responsiveness.
610 chars✓ has a “when” triggerlonger than Claude Code's old 250-char listing cap (fine on current versions)
Advanced

Key capabilities

  • Stitch nearby active enhancer peaks using a defined window
  • Exclude proximal-promoter signal from enhancer regions
  • Rank stitched regions by total signal intensity
  • Identify super-enhancers using the hockey-stick inflection point
  • Cross-reference super-enhancers with ENCODE dELS
  • Identify transcription factors encoded by super-enhancers for core regulatory circuitry

How it works

This skill processes ChIP-seq data to identify super-enhancers by stitching nearby active enhancer peaks, excluding promoter regions, ranking by signal, and applying a hockey-stick inflection point analysis.

Inputs & outputs

You give it
H3K27ac, MED1, or BRD4 ChIP-seq data (BAM files) and peak calls (GFF format)
You get back
Identified super-enhancer regions and associated ranking/inflection point data

When to use bio-chip-seq-super-enhancers

  • Identify super-enhancers for cell-identity genes
  • Compare super-enhancer landscapes between conditions
  • Analyze master transcription factor networks

About this skill

Version Compatibility

Reference examples tested with: ROSE (stjude/ROSE, 2018+), ROSE2 (linlabbcm/rose2, 2021+), LILY (BoevaLab/LILY, 2020+), HOMER 4.11+, samtools 1.19+, bedtools 2.31+, GenomicRanges 1.54+.

The original Young-lab ROSE is Python 2; ROSE2 (linlabbcm/rose2) and the stjude/ROSE fork are the Python-3 implementations with the same algorithm. For hg38 data use stjude/ROSE (python ROSE_main.py, whose genomeDict includes HG38); rose2's released genomeDict covers only HG18/HG19/MM8/MM9/MM10/RN4/RN6, so rose2 -g HG38 fails. LILY (Boeva 2017) is a refactored implementation with input-control background subtraction for low-quality H3K27ac data.

Super-Enhancer Calling

"Identify super-enhancers driving cell identity / cancer biology" -> Stitch nearby active enhancer peaks (H3K27ac, MED1, or BRD4) within a stitching window, exclude proximal-promoter signal, rank by total signal, find the hockey-stick inflection point where signal sharply increases, and classify all stitched regions above the inflection as super-enhancers.

  • CLI (stjude/ROSE, Python 3): python ROSE_main.py -g HG38 -i peaks.gff -r h3k27ac.bam -c input.bam -s 12500 -t 2500 -o rose_out/
  • CLI (HOMER): findPeaks tag_dir/ -style super -i input_tag_dir/
  • CLI (LILY): variant with input-control background subtraction
  • R (custom hockey-stick): rank enhancers by signal, find tangent-line inflection

The SE concept (Whyte 2013) is a thresholding heuristic on a continuous signal distribution (Pott & Lieb 2015 Nat Genet), not a categorical biological category. Genetic dissection of super-enhancers (Hay 2016; Moorthy 2017) shows constituent elements contribute unequally and many are individually dispensable/redundant; the "SE" label is a useful operational definition for BET-inhibitor responsiveness and cell-identity gene regulation, not an absolute biological property.

Marker Choice: H3K27ac vs MED1 vs BRD4

MarkerCapturesWhen to prefer
H3K27acActive regulatory elements broadlyMost widely available; standard for SE definition since Whyte 2013
MED1Mediator complex accumulation (the defining biology)Direct readout of SE; less common antibody; lower signal-to-noise
BRD4BET cofactor accumulationMost predictive of BET-inhibitor responsiveness; clinical relevance
H3K27ac + MED1 intersectionHigh-confidence SEGold standard if both available
dELS from ENCODE cCREsCell-type-agnostic distal enhancer registryCross-reference; not SE-specific by itself

Operational rule: H3K27ac for discovery; MED1 or BRD4 ChIP for functional / therapeutic claims. SE called on H3K27ac alone may not respond to BET inhibitors; SE called on BRD4 will.

Algorithmic Taxonomy

ToolMethodStrengthFails when
ROSE (Whyte 2013)Stitch within 12.5 kb, exclude ±2.5 kb of TSS, rank by signal, hockey-stick inflectionOriginal; widely cited; canonical referenceOriginal Young-lab code is Python 2; run the stjude/ROSE Py3 fork instead
ROSE2 (linlabbcm/rose2)Same algorithm, Python 3 portMaintained; pip-installable rose2 console commandReleased genomeDict has no HG38 (HG18/HG19/MM8/MM9/MM10/RN4/RN6 only) -> use stjude/ROSE for hg38
LILY (Boeva 2017)ROSE-like with input-control background subtractionWorks on lower-quality H3K27ac data; subtracts inputAdds complexity; less validated; specific to neuroblastoma/glioma in original paper
HOMER -style superNative ROSE-like in HOMER framework; stitching without TSS exclusionIntegrated with HOMER workflowDifferent stitching defaults; not directly comparable to ROSE counts
Custom hockey-stick (R)Generic rank-by-signal + tangent inflectionFlexible; works on any signal definitionReinvents algorithm; verify against ROSE on known dataset

Most papers use ROSE/ROSE2 with default stitching (12.5 kb) and TSS exclusion (2.5 kb). This is the de facto standard for cross-paper comparison. HOMER's -style super produces different counts and is not directly comparable.

Decision Tree: SE Calling Workflow

ScenarioRecommended pipeline
Standard SE discovery, H3K27ac availableROSE2 with default -s 12500 -t 2500; input control for subtraction
Predict BET-inhibitor responseBRD4 ChIP -> ROSE2 (or H3K27ac SE intersected with BRD4 peaks)
Compare SE between conditions (drug treatment)ROSE2 per condition + spike-in normalization (HDACi/BETi/EZH2i need ChIP-Rx)
Build core regulatory circuitryROSE2 + Saint-Andre 2016 algorithm: identify TFs encoded by SE that bind own SE + cross-bind other SE-encoded TFs
Low-quality H3K27ac (low FRiP)LILY with input subtraction
Compare with ENCODE dELS atlasROSE2 + intersect with ENCODE cCRE dELS BED
Differential SE between conditionsROSE2 per condition + signal-quantitative differential (DiffBind on SE regions)

ROSE / ROSE2 Workflow

Goal: Identify super-enhancers by stitching nearby active enhancer peaks within a stitching distance and ranking by total signal.

Approach: Convert peaks to GFF, exclude promoter-proximal peaks via -t (TSS exclusion window), stitch enhancers within -s (default 12.5 kb), rank by total H3K27ac (or MED1/BRD4) signal, find the hockey-stick inflection point, classify regions above as super-enhancers.

# Install ROSE2 (Python 3 port; unmaintained ROSE Py2 not recommended)
git clone https://github.com/linlabbcm/rose2.git
pip install ./rose2

# Convert peaks BED to GFF (ROSE requires GFF input)
awk 'BEGIN{OFS="\t"} {print $1,"peaks","enhancer",$2,$3,".",$6,".","ID="NR}' \
    peaks.narrowPeak > peaks.gff

# Filter promoter peaks before SE calling (within 2.5 kb of TSS)
# ROSE handles this via -t flag; preferable to pre-filter for clarity
bedtools intersect -a peaks.narrowPeak -b promoters_2kb.bed -v > enhancer_peaks.bed

# Run stjude/ROSE with input control (Python-3 fork; genomeDict includes HG38)
python ROSE_main.py -g HG38 -i peaks.gff \
    -r h3k27ac.bam -c input.bam \
    -o rose_output/ \
    -s 12500 \
    -t 2500

ROSE outputs:

  • *_AllEnhancers.table.txt — all stitched enhancer regions ranked by signal
  • *_SuperEnhancers.table.txt — SE only (above hockey-stick inflection)
  • *_Enhancers_withSuper.bed — BED with SE / TE classification
  • *_Plot_points.png — hockey-stick plot

Cross-Condition SE Comparison

This is the analysis most often done wrong. SE calling thresholds depend on absolute signal, so any global shift (HDACi, BETi, EZH2i) confounds direct SE-count comparison.

Wrong approach: Call ROSE2 on condition A and condition B separately, intersect SE BEDs, report "gained/lost SE."

Right approach:

  1. Spike-in normalize signal between conditions (see chip-seq/spike-in-normalization)
  2. Build a union SE set from both conditions
  3. Quantify signal at union SE regions per condition (DiffBind on the union)
  4. Apply differential testing with appropriate normalization (background-bin TMM or spike-in)
library(DiffBind)
# Union of SE BED files from condition A and B
union_se <- rtracklayer::import('union_SE.bed')
# Run DiffBind quantification on this region set with spike-in normalization

For BET-inhibitor experiments: the biology IS that all SE decrease globally; spike-in is mandatory.

Core Regulatory Circuitry (Saint-André 2016)

The CRC algorithm identifies master TF networks from SE annotations:

  1. List all TFs encoded by SE-associated genes
  2. For each such TF, check if its motif appears in its own SE (auto-regulation)
  3. Build a graph where TFs encoded by SE-A bind to motifs in SE-B
  4. Identify highly-interconnected sub-networks (CRC)
# Install CRC pipeline (console command is `crc`; -g is the genome BUILD, not a GTF)
pip install git+https://github.com/linlabcode/CRC.git

# Requires: SE enhancer table, subpeak BED, chromosome-FASTA dir
crc -e SE_table.txt -g HG38 -s subpeaks.bed -c chroms/ -o crc_out/ -n SAMPLE

CRC outputs the connected components of the regulatory network. Master TFs typically appear in the largest component with high out-degree.

ENCODE dELS Cross-Reference

ENCODE distal Enhancer-Like Signatures (dELS) are the cell-type-agnostic regulatory atlas (see chip-seq/peak-annotation). Cross-referencing SE against dELS:

  • Validates SE constituents are at canonical regulatory elements
  • Identifies SE constituents NOT in the dELS registry (potentially cell-type-specific)
  • Provides chromatin-state context (DNase + H3K27ac signatures)
wget https://downloads.wenglab.org/Registry-V4/GRCh38-cCREs.bed
awk -F'\t' '$NF == "dELS"' GRCh38-cCREs.bed > dels.bed

# Fraction of SE constituents overlapping dELS
bedtools intersect -a SuperEnhancers.bed -b dels.bed -u | wc -l
bedtools intersect -a SuperEnhancers.bed -b dels.bed -wa -wb > se_with_dels.tsv

Per-Tool Failure Modes

ROSE -- Python 2 dependency

Trigger: Running the original Young-lab ROSE_main.py (Python 2 code) on a modern system.

Mechanism: The original ROSE is Python 2 code; print statements without parens, dict.iteritems(), etc.

Symptom: SyntaxError on first import.

Fix: Use the stjude/ROSE fork (python ROSE_main.py, Python 3, genomeDict includes HG38) or ROSE2 (rose2, Python 3, but its released genomeDict has no HG38 -- HG18/HG19/MM8/MM9/MM10/RN4/RN6 only); identical algorithm and output format.

ROSE / ROSE2 -- Stitching distance default not appropriate for all biology

Trigger: Using default -s 12500 (12.5 kb) on small genomes or compact gene structures.

Mechanism: Default was set on human/mouse vertebrate genomes; Drosophila / yeast / plants have different regulatory architecture.

Fix: For non-vertebrate genomes, reduce stitching distance proportionally (e.g., -s 2500 for Drosophila, -s 500 for yeast).

ROSE / ROSE


Content truncated.

When not to use it

  • When the original Young-lab ROSE (Python 2) is used instead of Python 3 forks
  • When ROSE2's released genomeDict does not cover the target genome (e.g., HG38)
  • When HOMER's -style super is used for direct comparison with ROSE counts

Limitations

  • Original Young-lab ROSE is Python 2 and requires a Python 3 fork
  • ROSE2's released genomeDict does not include HG38
  • HOMER's -style super produces different counts not directly comparable to ROSE

How it compares

This skill automates the identification of super-enhancers from ChIP-seq data using established algorithms, providing a standardized and quantitative method compared to manual visual inspection of genomic regions.

Compared to similar skills

bio-chip-seq-super-enhancers side by side with the closest alternatives in the catalog.

SkillInstallsUpdatedSafetyDifficulty
bio-chip-seq-super-enhancers (this skill)128dReviewAdvanced
quant-analyst1032moNo flagsAdvanced
umap-learn62moReviewIntermediate
embedding-strategies82moNo flagsIntermediate

Try saying

Example prompts that trigger this skill in your AI assistant.

You might also like

quant-analyst

zenobi-us

Expert quantitative analyst specializing in financial modeling, algorithmic trading, and risk analytics. Masters statistical methods, derivatives pricing, and high-frequency trading with focus on mathematical rigor, performance optimization, and profitable strategy development.

103355

umap-learn

K-Dense-AI

UMAP dimensionality reduction. Fast nonlinear manifold learning for 2D/3D visualization, clustering preprocessing (HDBSCAN), supervised/parametric UMAP, for high-dimensional data.

6100

embedding-strategies

wshobson

Select and optimize embedding models for semantic search and RAG applications. Use when choosing embedding models, implementing chunking strategies, or optimizing embedding quality for specific domains.

890

building-automl-pipelines

jeremylongshore

Build automated machine learning pipelines, including feature engineering, model selection, and performance evaluation.

688

model-compare

rawwerks

Compare 3D CAD models using boolean operations (IoU, Dice, precision/recall). Use when evaluating generated models against gold references, diffing CAD revisions, or computing similarity metrics for ML training. Triggers on: model diff, compare models, IoU, intersection over union, model similarity, CAD comparison, STEP diff, 3D evaluation, gold reference, generated model, precision recall 3D.

783

matchms

davila7

Mass spectrometry analysis. Process mzML/MGF/MSP, spectral similarity (cosine, modified cosine), metadata harmonization, compound ID, for metabolomics and MS data processing.

674

Search skills

Search the agent skills registry