BU

bulk-rna-seq-batch-correction-with-combat

Harmonizes multi-batch RNA-seq data to remove technical variations before statistical analysis.

Install

mkdir -p .claude/skills/bulk-rna-seq-batch-correction-with-combat && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/2787" && unzip -o skill.zip -d .claude/skills/bulk-rna-seq-batch-correction-with-combat && rm skill.zip

Installs to .claude/skills/bulk-rna-seq-batch-correction-with-combat

Activation

This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.

Bulk RNA-seq batch correction with pyComBat: remove batch effects from merged cohorts, export corrected matrices, and benchmark visualizations.
143 charsno explicit “when” trigger
Intermediate

Key capabilities

  • Concatenates multi-cohort expression matrices
  • Maps genes to shared intersection across batches
  • Runs pyComBat batch correction
  • Generates pre/post correction visualization benchmarks

How it works

Uses the pyComBat wrapper to estimate and subtract additive/multiplicative batch effects from expression data.

Inputs & outputs

You give it
Multiple expression matrices and batch metadata
You get back
Harmonized AnnData object and corrected CSV matrices

When to use bulk-rna-seq-batch-correction-with-combat

  • Correct batch effects in merged bulk RNA-seq datasets
  • Harmonize microarray cohorts for downstream comparative analysis
  • Generate pre and post-correction plots for validation

About this skill

Bulk RNA-seq batch correction with ComBat

Overview

Apply this skill when a user has multiple bulk expression matrices measured across different batches and needs to harmonise them before downstream analysis. It follows t_bulk_combat.ipynb, w hich demonstrates the pyComBat workflow on ovarian cancer microarray cohorts.

Instructions

  1. Import core libraries
    • Load omicverse as ov, anndata, pandas as pd, and matplotlib.pyplot as plt.
    • Call ov.ov_plot_set() (aliased ov.plot_set() in some releases) to align figures with omicverse styling.
  2. Load each batch separately
    • Read the prepared pickled matrices (or user-provided expression tables) with pd.read_pickle(...)/pd.read_csv(...).
    • Transpose to gene × sample before wrapping them in anndata.AnnData objects so adata.obs stores sample metadata.
    • Assign a batch column for every cohort (adata.obs['batch'] = '1', '2', ...). Encourage descriptive labels when availa ble.
  3. Concatenate on shared genes
    • Use anndata.concat([adata1, adata2, adata3], merge='same') to retain the intersection of genes across batches.
    • Confirm the combined adata reports balanced sample counts per batch; if not, prompt users to re-check inputs.
  4. Run ComBat batch correction
    • Execute ov.bulk.batch_correction(adata, batch_key='batch').
    • Explain that corrected values are stored in adata.layers['batch_correction'] while the original counts remain in adata.X.
  5. Export corrected and raw matrices
    • Obtain DataFrames via adata.to_df().T (raw) and adata.to_df(layer='batch_correction').T (corrected).
    • Encourage saving both tables (.to_csv(...)) plus the harmonised AnnData (adata.write_h5ad('adata_batch.h5ad', compressio n='gzip')).
  6. Benchmark the correction
    • For per-sample variance checks, draw before/after boxplots and recolour boxes using ov.pl.red_color, blue_color, gree n_color palettes to match batches.
    • Copy raw counts to a named layer with adata.layers['raw'] = adata.X.copy() before PCA.
    • Run ov.pp.pca(adata, layer='raw', n_pcs=50) and ov.pp.pca(adata, layer='batch_correction', n_pcs=50).
    • Visualise embeddings with ov.pl.embedding(..., basis='raw|original|X_pca', color='batch', frameon='small') and repeat fo r the corrected layer to verify mixing.
  7. Defensive validation
    # Before ComBat: verify batch column exists and has >1 batch
    assert 'batch' in adata.obs.columns, "adata.obs must contain a 'batch' column"
    n_batches = adata.obs['batch'].nunique()
    assert n_batches > 1, f"Only {n_batches} batch — need >1 for batch correction"
    # Verify gene overlap after concatenation
    if adata.n_vars < 100:
        print(f"WARNING: Only {adata.n_vars} shared genes after concat — check gene ID harmonization")
    
  8. Troubleshooting tips
    • Mismatched gene identifiers cause dropped features—remind users to harmonise feature names (e.g., gene symbols) before conca tenation.
    • pyComBat expects log-scale intensities or similarly distributed counts; recommend log-transforming strongly skewed matrices.
    • If batch_correction layer is missing, ensure the batch_key matches the column name in adata.obs.

Examples

  • "Combine three GEO ovarian cohorts, run ComBat, and export both the raw and corrected CSV matrices."
  • "Plot PCA embeddings before and after batch correction to confirm that batches 1–3 overlap."
  • "Save the harmonised AnnData file so I can reload it later for downstream DEG analysis."

References

When not to use it

  • Datasets with no shared gene markers
  • Single-batch expression analysis

Prerequisites

pythonomicversepandasanndata

Limitations

  • Requires balanced sample counts for optimal results
  • Cannot recover data lost during gene filtering
  • Assumes batch effects are the primary signal bias

How it compares

It standardizes the harmonization process across disparate cohorts rather than performing manual adjustment.

Compared to similar skills

bulk-rna-seq-batch-correction-with-combat side by side with the closest alternatives in the catalog.

SkillInstallsUpdatedSafetyDifficulty
bulk-rna-seq-batch-correction-with-combat (this skill)15moNo flagsIntermediate
quant-analyst1032moNo flagsAdvanced
umap-learn62moReviewIntermediate
embedding-strategies82moNo flagsIntermediate

Try saying

Example prompts that trigger this skill in your AI assistant.

You might also like

quant-analyst

zenobi-us

Expert quantitative analyst specializing in financial modeling, algorithmic trading, and risk analytics. Masters statistical methods, derivatives pricing, and high-frequency trading with focus on mathematical rigor, performance optimization, and profitable strategy development.

103355

umap-learn

K-Dense-AI

UMAP dimensionality reduction. Fast nonlinear manifold learning for 2D/3D visualization, clustering preprocessing (HDBSCAN), supervised/parametric UMAP, for high-dimensional data.

6100

embedding-strategies

wshobson

Select and optimize embedding models for semantic search and RAG applications. Use when choosing embedding models, implementing chunking strategies, or optimizing embedding quality for specific domains.

890

building-automl-pipelines

jeremylongshore

Build automated machine learning pipelines, including feature engineering, model selection, and performance evaluation.

688

model-compare

rawwerks

Compare 3D CAD models using boolean operations (IoU, Dice, precision/recall). Use when evaluating generated models against gold references, diffing CAD revisions, or computing similarity metrics for ML training. Triggers on: model diff, compare models, IoU, intersection over union, model similarity, CAD comparison, STEP diff, 3D evaluation, gold reference, generated model, precision recall 3D.

783

matchms

davila7

Mass spectrometry analysis. Process mzML/MGF/MSP, spectral similarity (cosine, modified cosine), metadata harmonization, compound ID, for metabolomics and MS data processing.

674

Search skills

Search the agent skills registry