BU

bulktrajblend-trajectory-interpolation

Integrates bulk and single-cell RNA-seq to interpolate gaps in developmental trajectories.

Install

mkdir -p .claude/skills/bulktrajblend-trajectory-interpolation && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/2785" && unzip -o skill.zip -d .claude/skills/bulktrajblend-trajectory-interpolation && rm skill.zip

Installs to .claude/skills/bulktrajblend-trajectory-interpolation

Activation

This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.

Extend scRNA-seq developmental trajectories with BulkTrajBlend by generating intermediate cells from bulk RNA-seq, training beta-VAE and GNN models, and interpolating missing states.
182 charsno explicit “when” trigger
Advanced

Key capabilities

  • Deconvolves bulk RNA-seq samples
  • Trains beta-VAE generative models
  • Interpolates missing developmental states
  • Integrates bulk and single-cell data

How it works

Trains a VAE to learn the latent space of cell states, allowing for the generation of synthetic data points based on bulk input.

Inputs & outputs

You give it
Bulk expression counts and scRNA-seq AnnData
You get back
Augmented AnnData with synthetic intermediate cells

When to use bulktrajblend-trajectory-interpolation

  • Extend incomplete scRNA-seq trajectories using bulk data
  • Deconvolve bulk tissue samples into cell types
  • Train VAE models to interpolate developmental stages

About this skill

BulkTrajBlend trajectory interpolation

Overview

Invoke this skill when users need to bridge gaps in single-cell developmental trajectories using matched bulk RNA-seq. It follows t_bulktrajblend.ipynb, showcasing how BulkTrajBlend deconvolves PDAC bulk samples, identifies overlapping communities with a GNN, and interpolates "interrupted" cell states.

Instructions

  1. Prepare libraries and inputs
    • Import omicverse as ov, scanpy as sc, scvelo as scv, and helper functions like from omicverse.utils import mde; run ov.plot_set().
    • Load the reference scRNA-seq AnnData (scv.datasets.dentategyrus()) and raw bulk counts with ov.utils.read(...) followed by ov.bulk.Matrix_ID_mapping(...) for gene ID harmonisation.
  2. Configure BulkTrajBlend
    • Instantiate ov.bulk2single.BulkTrajBlend(bulk_seq=bulk_df, single_seq=adata, bulk_group=['dg_d_1','dg_d_2','dg_d_3'], celltype_key='clusters').
    • Explain that bulk_group names correspond to raw bulk columns and the method expects unscaled counts.
  3. Set beta-VAE expectations
    • Call bulktb.vae_configure(cell_target_num=100) (or pass a dictionary) to define expected cell counts per cluster. Mention that omitting the argument triggers TAPE-based estimation.
  4. Train or load the beta-VAE
    • Use bulktb.vae_train(batch_size=512, learning_rate=1e-4, hidden_size=256, epoch_num=3500, vae_save_dir='...', vae_save_name='dg_btb_vae', generate_save_dir='...', generate_save_name='dg_btb').
    • Highlight resuming with bulktb.vae_load('.../dg_btb_vae.pth') and the need to regenerate cells with consistent random seeds for reproducibility.
  5. Generate synthetic cells
    • Produce filtered AnnData via bulktb.vae_generate(leiden_size=25) and inspect compositions with ov.bulk2single.bulk2single_plot_cellprop(...).
    • Save outputs to disk for reuse (adata.write_h5ad).
  6. Configure and train the GNN
    • Call bulktb.gnn_configure(max_epochs=2000, use_rep='X', neighbor_rep='X_pca', gpu=0, ...) to set hyperparameters.
    • Train using bulktb.gnn_train(); reload checkpoints with bulktb.gnn_load('save_model/gnn.pth').
    • Generate overlapping community assignments through bulktb.gnn_generate().
  7. Visualise community structure
    • Create MDE embeddings: bulktb.nocd_obj.adata.obsm['X_mde'] = mde(bulktb.nocd_obj.adata.obsm['X_pca']).
    • Plot clusters vs. discovered communities using sc.pl.embedding(..., color=['clusters','nocd_n'], palette=ov.utils.pyomic_palette()) and filtered subsets excluding synthetic labels with hyphens.
  8. Interpolate missing states
    • Run bulktb.interpolation('OPC') (replace with target lineage) to synthesise continuity, then preprocess the interpolated AnnData (HVG selection, scaling, PCA).
    • Compute embeddings with mde, visualise with ov.pl.embedding, and compare to the original atlas.
  9. Analyse trajectories
    • Initialise ov.single.pyVIA on both original and interpolated data to derive pseudotime, followed by get_pseudotime, ov.pp.neighbors, ov.utils.cal_paga, and ov.utils.plot_paga for topology validation.
  10. Defensive validation
    # Before BulkTrajBlend: verify bulk_group columns exist
    for g in bulk_group:
        assert g in bulk_df.columns, f"Bulk group '{g}' not in bulk data columns"
    # Verify celltype_key exists in reference
    assert celltype_key in adata.obs.columns, f"Cell type column '{celltype_key}' not in reference AnnData"
    # Verify gene name overlap
    shared = set(bulk_df.index) & set(adata.var_names)
    assert len(shared) > 100, f"Only {len(shared)} shared genes — harmonize gene IDs first"
    
  11. Troubleshooting tips
    • If the VAE collapses (high reconstruction loss), lower learning_rate or reduce hidden_size.
    • Ensure the same generated dataset is used before calling gnn_train; regenerating cells changes the graph and can break checkpoint loading.
    • Sparse clusters may need adjusted cell_target_num thresholds or a smaller leiden_size filter to retain rare populations.

Examples

  • "Train BulkTrajBlend on PDAC cohorts, then interpolate missing OPC states in the trajectory."
  • "Load saved beta-VAE and GNN weights to regenerate overlapping communities and plot cluster vs. nocd labels."
  • "Run VIA on interpolated cells and compare PAGA graphs with the original scRNA-seq trajectory."

References

When not to use it

  • Datasets with no matched bulk reference
  • Small-scale experimental data

Prerequisites

omicversescanpyscvelo

Limitations

  • High computational cost for VAE training
  • Requires consistent random seeds for results
  • Model performance depends on cell state overlap

How it compares

It extends limited single-cell trajectories using bulk tissue signatures instead of relying solely on sparse single-cell data.

Compared to similar skills

bulktrajblend-trajectory-interpolation side by side with the closest alternatives in the catalog.

SkillInstallsUpdatedSafetyDifficulty
bulktrajblend-trajectory-interpolation (this skill)35moNo flagsAdvanced
quant-analyst1032moNo flagsAdvanced
umap-learn62moReviewIntermediate
embedding-strategies82moNo flagsIntermediate

Try saying

Example prompts that trigger this skill in your AI assistant.

You might also like

quant-analyst

zenobi-us

Expert quantitative analyst specializing in financial modeling, algorithmic trading, and risk analytics. Masters statistical methods, derivatives pricing, and high-frequency trading with focus on mathematical rigor, performance optimization, and profitable strategy development.

103355

umap-learn

K-Dense-AI

UMAP dimensionality reduction. Fast nonlinear manifold learning for 2D/3D visualization, clustering preprocessing (HDBSCAN), supervised/parametric UMAP, for high-dimensional data.

6100

embedding-strategies

wshobson

Select and optimize embedding models for semantic search and RAG applications. Use when choosing embedding models, implementing chunking strategies, or optimizing embedding quality for specific domains.

890

building-automl-pipelines

jeremylongshore

Build automated machine learning pipelines, including feature engineering, model selection, and performance evaluation.

688

model-compare

rawwerks

Compare 3D CAD models using boolean operations (IoU, Dice, precision/recall). Use when evaluating generated models against gold references, diffing CAD revisions, or computing similarity metrics for ML training. Triggers on: model diff, compare models, IoU, intersection over union, model similarity, CAD comparison, STEP diff, 3D evaluation, gold reference, generated model, precision recall 3D.

783

matchms

davila7

Mass spectrometry analysis. Process mzML/MGF/MSP, spectral similarity (cosine, modified cosine), metadata harmonization, compound ID, for metabolomics and MS data processing.

674

Search skills

Search the agent skills registry