bio-chip-seq-super-enhancers
Detects super-enhancers from H3K27ac/MED1/BRD4 ChIP-seq data to study cell-identity gene regulation.
Install
mkdir -p .claude/skills/bio-chip-seq-super-enhancers && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/5814" && unzip -o skill.zip -d .claude/skills/bio-chip-seq-super-enhancers && rm skill.zipInstalls to .claude/skills/bio-chip-seq-super-enhancers
Activation
This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.
Identifies super-enhancers from H3K27ac, MED1, or BRD4 ChIP-seq using ROSE, ROSE2, LILY, HOMER -style super, and ENCODE dELS cross-referencing. Handles peak stitching parameters, ranking choices, hockey-stick inflection, marker choice (H3K27ac vs MED1/BRD4), and cross-condition comparison with spike-in normalization. Constructs core regulatory circuitry (Saint-Andre 2016) from SE-encoded TFs. Use when identifying cell-identity / cancer-associated regulatory domains, comparing super-enhancers between conditions, identifying master transcription factor networks, or predicting BET-inhibitor responsiveness.Key capabilities
- →Stitch nearby active enhancer peaks using a defined window
- →Exclude proximal-promoter signal from enhancer regions
- →Rank stitched regions by total signal intensity
- →Identify super-enhancers using the hockey-stick inflection point
- →Cross-reference super-enhancers with ENCODE dELS
- →Identify transcription factors encoded by super-enhancers for core regulatory circuitry
How it works
This skill processes ChIP-seq data to identify super-enhancers by stitching nearby active enhancer peaks, excluding promoter regions, ranking by signal, and applying a hockey-stick inflection point analysis.
Inputs & outputs
When to use bio-chip-seq-super-enhancers
- →Identify super-enhancers for cell-identity genes
- →Compare super-enhancer landscapes between conditions
- →Analyze master transcription factor networks
About this skill
Version Compatibility
Reference examples tested with: ROSE (stjude/ROSE, 2018+), ROSE2 (linlabbcm/rose2, 2021+), LILY (BoevaLab/LILY, 2020+), HOMER 4.11+, samtools 1.19+, bedtools 2.31+, GenomicRanges 1.54+.
The original Young-lab ROSE is Python 2; ROSE2 (linlabbcm/rose2) and the stjude/ROSE fork are the Python-3 implementations with the same algorithm. For hg38 data use stjude/ROSE (python ROSE_main.py, whose genomeDict includes HG38); rose2's released genomeDict covers only HG18/HG19/MM8/MM9/MM10/RN4/RN6, so rose2 -g HG38 fails. LILY (Boeva 2017) is a refactored implementation with input-control background subtraction for low-quality H3K27ac data.
Super-Enhancer Calling
"Identify super-enhancers driving cell identity / cancer biology" -> Stitch nearby active enhancer peaks (H3K27ac, MED1, or BRD4) within a stitching window, exclude proximal-promoter signal, rank by total signal, find the hockey-stick inflection point where signal sharply increases, and classify all stitched regions above the inflection as super-enhancers.
- CLI (stjude/ROSE, Python 3):
python ROSE_main.py -g HG38 -i peaks.gff -r h3k27ac.bam -c input.bam -s 12500 -t 2500 -o rose_out/ - CLI (HOMER):
findPeaks tag_dir/ -style super -i input_tag_dir/ - CLI (LILY): variant with input-control background subtraction
- R (custom hockey-stick): rank enhancers by signal, find tangent-line inflection
The SE concept (Whyte 2013) is a thresholding heuristic on a continuous signal distribution (Pott & Lieb 2015 Nat Genet), not a categorical biological category. Genetic dissection of super-enhancers (Hay 2016; Moorthy 2017) shows constituent elements contribute unequally and many are individually dispensable/redundant; the "SE" label is a useful operational definition for BET-inhibitor responsiveness and cell-identity gene regulation, not an absolute biological property.
Marker Choice: H3K27ac vs MED1 vs BRD4
| Marker | Captures | When to prefer |
|---|---|---|
| H3K27ac | Active regulatory elements broadly | Most widely available; standard for SE definition since Whyte 2013 |
| MED1 | Mediator complex accumulation (the defining biology) | Direct readout of SE; less common antibody; lower signal-to-noise |
| BRD4 | BET cofactor accumulation | Most predictive of BET-inhibitor responsiveness; clinical relevance |
| H3K27ac + MED1 intersection | High-confidence SE | Gold standard if both available |
| dELS from ENCODE cCREs | Cell-type-agnostic distal enhancer registry | Cross-reference; not SE-specific by itself |
Operational rule: H3K27ac for discovery; MED1 or BRD4 ChIP for functional / therapeutic claims. SE called on H3K27ac alone may not respond to BET inhibitors; SE called on BRD4 will.
Algorithmic Taxonomy
| Tool | Method | Strength | Fails when |
|---|---|---|---|
| ROSE (Whyte 2013) | Stitch within 12.5 kb, exclude ±2.5 kb of TSS, rank by signal, hockey-stick inflection | Original; widely cited; canonical reference | Original Young-lab code is Python 2; run the stjude/ROSE Py3 fork instead |
| ROSE2 (linlabbcm/rose2) | Same algorithm, Python 3 port | Maintained; pip-installable rose2 console command | Released genomeDict has no HG38 (HG18/HG19/MM8/MM9/MM10/RN4/RN6 only) -> use stjude/ROSE for hg38 |
| LILY (Boeva 2017) | ROSE-like with input-control background subtraction | Works on lower-quality H3K27ac data; subtracts input | Adds complexity; less validated; specific to neuroblastoma/glioma in original paper |
HOMER -style super | Native ROSE-like in HOMER framework; stitching without TSS exclusion | Integrated with HOMER workflow | Different stitching defaults; not directly comparable to ROSE counts |
| Custom hockey-stick (R) | Generic rank-by-signal + tangent inflection | Flexible; works on any signal definition | Reinvents algorithm; verify against ROSE on known dataset |
Most papers use ROSE/ROSE2 with default stitching (12.5 kb) and TSS exclusion (2.5 kb). This is the de facto standard for cross-paper comparison. HOMER's -style super produces different counts and is not directly comparable.
Decision Tree: SE Calling Workflow
| Scenario | Recommended pipeline |
|---|---|
| Standard SE discovery, H3K27ac available | ROSE2 with default -s 12500 -t 2500; input control for subtraction |
| Predict BET-inhibitor response | BRD4 ChIP -> ROSE2 (or H3K27ac SE intersected with BRD4 peaks) |
| Compare SE between conditions (drug treatment) | ROSE2 per condition + spike-in normalization (HDACi/BETi/EZH2i need ChIP-Rx) |
| Build core regulatory circuitry | ROSE2 + Saint-Andre 2016 algorithm: identify TFs encoded by SE that bind own SE + cross-bind other SE-encoded TFs |
| Low-quality H3K27ac (low FRiP) | LILY with input subtraction |
| Compare with ENCODE dELS atlas | ROSE2 + intersect with ENCODE cCRE dELS BED |
| Differential SE between conditions | ROSE2 per condition + signal-quantitative differential (DiffBind on SE regions) |
ROSE / ROSE2 Workflow
Goal: Identify super-enhancers by stitching nearby active enhancer peaks within a stitching distance and ranking by total signal.
Approach: Convert peaks to GFF, exclude promoter-proximal peaks via -t (TSS exclusion window), stitch enhancers within -s (default 12.5 kb), rank by total H3K27ac (or MED1/BRD4) signal, find the hockey-stick inflection point, classify regions above as super-enhancers.
# Install ROSE2 (Python 3 port; unmaintained ROSE Py2 not recommended)
git clone https://github.com/linlabbcm/rose2.git
pip install ./rose2
# Convert peaks BED to GFF (ROSE requires GFF input)
awk 'BEGIN{OFS="\t"} {print $1,"peaks","enhancer",$2,$3,".",$6,".","ID="NR}' \
peaks.narrowPeak > peaks.gff
# Filter promoter peaks before SE calling (within 2.5 kb of TSS)
# ROSE handles this via -t flag; preferable to pre-filter for clarity
bedtools intersect -a peaks.narrowPeak -b promoters_2kb.bed -v > enhancer_peaks.bed
# Run stjude/ROSE with input control (Python-3 fork; genomeDict includes HG38)
python ROSE_main.py -g HG38 -i peaks.gff \
-r h3k27ac.bam -c input.bam \
-o rose_output/ \
-s 12500 \
-t 2500
ROSE outputs:
*_AllEnhancers.table.txt— all stitched enhancer regions ranked by signal*_SuperEnhancers.table.txt— SE only (above hockey-stick inflection)*_Enhancers_withSuper.bed— BED with SE / TE classification*_Plot_points.png— hockey-stick plot
Cross-Condition SE Comparison
This is the analysis most often done wrong. SE calling thresholds depend on absolute signal, so any global shift (HDACi, BETi, EZH2i) confounds direct SE-count comparison.
Wrong approach: Call ROSE2 on condition A and condition B separately, intersect SE BEDs, report "gained/lost SE."
Right approach:
- Spike-in normalize signal between conditions (see chip-seq/spike-in-normalization)
- Build a union SE set from both conditions
- Quantify signal at union SE regions per condition (DiffBind on the union)
- Apply differential testing with appropriate normalization (background-bin TMM or spike-in)
library(DiffBind)
# Union of SE BED files from condition A and B
union_se <- rtracklayer::import('union_SE.bed')
# Run DiffBind quantification on this region set with spike-in normalization
For BET-inhibitor experiments: the biology IS that all SE decrease globally; spike-in is mandatory.
Core Regulatory Circuitry (Saint-André 2016)
The CRC algorithm identifies master TF networks from SE annotations:
- List all TFs encoded by SE-associated genes
- For each such TF, check if its motif appears in its own SE (auto-regulation)
- Build a graph where TFs encoded by SE-A bind to motifs in SE-B
- Identify highly-interconnected sub-networks (CRC)
# Install CRC pipeline (console command is `crc`; -g is the genome BUILD, not a GTF)
pip install git+https://github.com/linlabcode/CRC.git
# Requires: SE enhancer table, subpeak BED, chromosome-FASTA dir
crc -e SE_table.txt -g HG38 -s subpeaks.bed -c chroms/ -o crc_out/ -n SAMPLE
CRC outputs the connected components of the regulatory network. Master TFs typically appear in the largest component with high out-degree.
ENCODE dELS Cross-Reference
ENCODE distal Enhancer-Like Signatures (dELS) are the cell-type-agnostic regulatory atlas (see chip-seq/peak-annotation). Cross-referencing SE against dELS:
- Validates SE constituents are at canonical regulatory elements
- Identifies SE constituents NOT in the dELS registry (potentially cell-type-specific)
- Provides chromatin-state context (DNase + H3K27ac signatures)
wget https://downloads.wenglab.org/Registry-V4/GRCh38-cCREs.bed
awk -F'\t' '$NF == "dELS"' GRCh38-cCREs.bed > dels.bed
# Fraction of SE constituents overlapping dELS
bedtools intersect -a SuperEnhancers.bed -b dels.bed -u | wc -l
bedtools intersect -a SuperEnhancers.bed -b dels.bed -wa -wb > se_with_dels.tsv
Per-Tool Failure Modes
ROSE -- Python 2 dependency
Trigger: Running the original Young-lab ROSE_main.py (Python 2 code) on a modern system.
Mechanism: The original ROSE is Python 2 code; print statements without parens, dict.iteritems(), etc.
Symptom: SyntaxError on first import.
Fix: Use the stjude/ROSE fork (python ROSE_main.py, Python 3, genomeDict includes HG38) or ROSE2 (rose2, Python 3, but its released genomeDict has no HG38 -- HG18/HG19/MM8/MM9/MM10/RN4/RN6 only); identical algorithm and output format.
ROSE / ROSE2 -- Stitching distance default not appropriate for all biology
Trigger: Using default -s 12500 (12.5 kb) on small genomes or compact gene structures.
Mechanism: Default was set on human/mouse vertebrate genomes; Drosophila / yeast / plants have different regulatory architecture.
Fix: For non-vertebrate genomes, reduce stitching distance proportionally (e.g., -s 2500 for Drosophila, -s 500 for yeast).
ROSE / ROSE
Content truncated.
When not to use it
- →When the original Young-lab ROSE (Python 2) is used instead of Python 3 forks
- →When ROSE2's released genomeDict does not cover the target genome (e.g., HG38)
- →When HOMER's -style super is used for direct comparison with ROSE counts
Limitations
- →Original Young-lab ROSE is Python 2 and requires a Python 3 fork
- →ROSE2's released genomeDict does not include HG38
- →HOMER's -style super produces different counts not directly comparable to ROSE
How it compares
This skill automates the identification of super-enhancers from ChIP-seq data using established algorithms, providing a standardized and quantitative method compared to manual visual inspection of genomic regions.
Compared to similar skills
bio-chip-seq-super-enhancers side by side with the closest alternatives in the catalog.
| Skill | Installs | Updated | Safety | Difficulty |
|---|---|---|---|---|
| bio-chip-seq-super-enhancers (this skill) | 1 | 28d | Review | Advanced |
| quant-analyst | 103 | 2mo | No flags | Advanced |
| umap-learn | 6 | 2mo | Review | Intermediate |
| embedding-strategies | 8 | 2mo | No flags | Intermediate |
Try saying
Example prompts that trigger this skill in your AI assistant.
More by GPTomics
View all by GPTomics →You might also like
quant-analyst
zenobi-us
Expert quantitative analyst specializing in financial modeling, algorithmic trading, and risk analytics. Masters statistical methods, derivatives pricing, and high-frequency trading with focus on mathematical rigor, performance optimization, and profitable strategy development.
umap-learn
K-Dense-AI
UMAP dimensionality reduction. Fast nonlinear manifold learning for 2D/3D visualization, clustering preprocessing (HDBSCAN), supervised/parametric UMAP, for high-dimensional data.
embedding-strategies
wshobson
Select and optimize embedding models for semantic search and RAG applications. Use when choosing embedding models, implementing chunking strategies, or optimizing embedding quality for specific domains.
building-automl-pipelines
jeremylongshore
Build automated machine learning pipelines, including feature engineering, model selection, and performance evaluation.
model-compare
rawwerks
Compare 3D CAD models using boolean operations (IoU, Dice, precision/recall). Use when evaluating generated models against gold references, diffing CAD revisions, or computing similarity metrics for ML training. Triggers on: model diff, compare models, IoU, intersection over union, model similarity, CAD comparison, STEP diff, 3D evaluation, gold reference, generated model, precision recall 3D.
matchms
davila7
Mass spectrometry analysis. Process mzML/MGF/MSP, spectral similarity (cosine, modified cosine), metadata harmonization, compound ID, for metabolomics and MS data processing.