BI

bio-alignment-indexing

Generates genomic alignment indices for fast data retrieval.

Install

mkdir -p .claude/skills/bio-alignment-indexing && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/9349" && unzip -o skill.zip -d .claude/skills/bio-alignment-indexing && rm skill.zip

Installs to .claude/skills/bio-alignment-indexing

Activation

This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.

Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam. Use when enabling random access to alignment files or fetching specific genomic regions.
164 chars✓ has a “when” trigger
Intermediate

Key capabilities

  • Generates.bai and.csi indices for alignment files
  • Selects appropriate index type based on contig size
  • Configures min_shift parameters for CSI indices
  • Provides Python code using pysam for indexing
  • Validates version compatibility for samtools

How it works

It calculates the bin requirements for the alignment file's largest contig and selects the index type (BAI vs CSI) based on thresholds. It verifies tool compatibility before recommending specific CLI/API flags.

Inputs & outputs

You give it
BAM/CRAM file details and reference genome stats
You get back
Indexing command or Python code snippet

When to use bio-alignment-indexing

  • Index BAM files for random access
  • Generate CSI index for large contigs
  • Fetch specific genomic regions from alignment

About this skill

Version Compatibility

Reference examples tested with: pysam 0.22+, samtools 1.19+

Before using code patterns, verify installed versions match. If versions differ:

  • Python: pip show <package> then help(module.function) to check signatures
  • CLI: <tool> --version then <tool> --help to confirm flags

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

Alignment Indexing

Create indices for random access to alignment files using samtools and pysam.

"Index a BAM file" -> Create a .bai/.csi index enabling random access to genomic regions.

  • CLI: samtools index file.bam
  • Python: pysam.index('file.bam')

Index Types

IndexExtensionMax contigBin shiftWhen required
BAI.bai / .bam.bai2^29 bp ≈ 537 Mbpfixed (16 kb)Default for human, mouse, fly, fish
CSI.csi / .bam.csi2^(min_shift + depth*3)configurable via -mRequired for any contig >537 Mbp
CRAI.crai / .cram.craichunk-basedn/aCRAM only
TBI.tbi2^29-1fixedtabix VCF/BED -- same limit as BAI

Which Index for Which Genome

GenomeLargest contigIndex
GRCh38 / GRCh37 (human)248 MbpBAI
GRCm39 (mouse)195 MbpBAI
GRCz11 (zebrafish), TAIR10 (Arabidopsis)78 Mbp / 30 MbpBAI
Wheat IWGSC (Triticum aestivum)~830 Mbp (chr3B)CSI
Pine, fir, axolotl, sugar pinemulti-GbpCSI with larger -m
Long-read assembly with very large contigsvariescheck cut -f2 ref.fa.fai | sort -nr | head -1

For polyploid plants and salamander-scale genomes, increase the bin shift:

# Default CSI matches BAI bin layout: 2^(14 + 5*3) = 2^29 bp ≈ 537 Mbp per contig
samtools index -c file.bam

# Larger min_shift for contigs >537 Mbp (wheat, axolotl, sugar pine)
samtools index -c -m 18 file.bam   # 2^(18+15) = 2^33 = ~8.5 Gbp per contig

Index file precedence: htslib (and therefore the samtools CLI) tries .csi before .bai on auto-load, so when both exist the .csi is used -- a stale .bai left over after re-indexing to CSI is generally ignored. (Note: htsjdk/Java prefers .bai, the opposite order.) Removing the obsolete .bai still avoids confusion.

samtools index

Create BAI Index

samtools index input.bam
# Creates input.bam.bai

Create CSI Index

samtools index -c input.bam
# Creates input.bam.csi

Specify Output Name

samtools index input.bam output.bai

Multi-threaded Indexing

samtools index -@ 4 input.bam

Index CRAM

samtools index input.cram
# Creates input.cram.crai

Index Requirements

Indexing requires coordinate-sorted files:

# Check sort order
samtools view -H input.bam | grep "^@HD"
# Should show SO:coordinate

# Sort if needed, then index
samtools sort -o sorted.bam input.bam
samtools index sorted.bam

Using Indices for Region Access

Goal: Extract reads overlapping specific genomic coordinates from an indexed BAM.

Approach: With the index present, samtools view or pysam.fetch() can jump directly to the relevant file offset instead of scanning the entire file.

samtools view with Region

# Requires index file present
samtools view input.bam chr1:1000000-2000000

Multiple Regions

samtools view input.bam chr1:1000-2000 chr2:3000-4000

Regions from BED File

samtools view -L regions.bed input.bam

pysam Python Alternative

Create Index

import pysam

pysam.index('input.bam')
# Creates input.bam.bai

Create CSI Index

# pysam.index passes through to samtools index; pass the -c flag for CSI.
pysam.index('-c', 'input.bam')
# Produces input.bam.csi.

Fetch with Index

with pysam.AlignmentFile('input.bam', 'rb') as bam:
    # fetch() requires index
    for read in bam.fetch('chr1', 1000000, 2000000):
        print(read.query_name)

Check if Indexed

import pysam
from pathlib import Path

def is_indexed(bam_path):
    bam_path = Path(bam_path)
    return (bam_path.with_suffix('.bam.bai').exists() or
            Path(str(bam_path) + '.bai').exists() or
            bam_path.with_suffix('.bam.csi').exists())

if not is_indexed('input.bam'):
    pysam.index('input.bam')

Fetch Multiple Regions

regions = [('chr1', 1000, 2000), ('chr1', 5000, 6000), ('chr2', 1000, 2000)]

with pysam.AlignmentFile('input.bam', 'rb') as bam:
    for chrom, start, end in regions:
        count = sum(1 for _ in bam.fetch(chrom, start, end))
        print(f'{chrom}:{start}-{end}: {count} reads')

Count Reads in Region

with pysam.AlignmentFile('input.bam', 'rb') as bam:
    count = bam.count('chr1', 1000000, 2000000)
    print(f'Reads in region: {count}')

Get Reads Covering Position

with pysam.AlignmentFile('input.bam', 'rb') as bam:
    for read in bam.fetch('chr1', 1000000, 1000001):
        if read.reference_start <= 1000000 < read.reference_end:
            print(f'{read.query_name} covers position 1000000')

Index File Locations

samtools looks for indices in two locations:

input.bam.bai   # Standard location
input.bai       # Alternative location

For CRAM:

input.cram.crai

idxstats - Index Statistics

Get Per-Chromosome Counts

samtools idxstats input.bam

Output format:

chr1    248956422    5000000    0
chr2    242193529    4500000    0
*       0            0          10000

Columns: reference name, length, mapped reads, unmapped reads.

What "mapped" Actually Counts (Caveat)

The mapped column counts every alignment record with that RNAME, including secondary AND supplementary. For long-read minimap2 output, where a single read can produce many supplementary chimeric alignments, idxstats overcounts input reads -- typically 1.5-3x.

For unique read counts, use primary-only:

samtools view -c -F 2304 input.bam chr1   # primary only

Cross-check unmapped consistency (a senior sanity check):

samtools idxstats file.bam | awk '{sum+=$4} END {print sum}'   # idxstats unmapped (sum across all rows; PE orphans get a contig RNAME)
samtools view -c -f 4 -F 2304 file.bam                         # primary unmapped (should match)

Sum Total Mapped Reads

samtools idxstats input.bam | awk '{sum += $3} END {print sum}'

pysam idxstats

with pysam.AlignmentFile('input.bam', 'rb') as bam:
    for stat in bam.get_index_statistics():
        print(f'{stat.contig}: {stat.mapped} mapped, {stat.unmapped} unmapped')

FASTA Index (faidx)

Related but different - index reference FASTA for random access:

samtools faidx reference.fa
# Creates reference.fa.fai

# Fetch region from indexed FASTA
samtools faidx reference.fa chr1:1000-2000

pysam FastaFile

with pysam.FastaFile('reference.fa') as ref:
    seq = ref.fetch('chr1', 1000, 2000)
    print(seq)

Quick Reference

Tasksamtoolspysam
Create BAIsamtools index file.bampysam.index('file.bam')
Create CSIsamtools index -c file.bampysam.index('-c', 'file.bam')
Fetch regionsamtools view file.bam chr1:1-1000bam.fetch('chr1', 0, 1000)
Count in regionsamtools view -c file.bam chr1:1-1000bam.count('chr1', 0, 1000)
Index statssamtools idxstats file.bambam.get_index_statistics()
Index FASTAsamtools faidx ref.faAutomatic with FastaFile

Index Staleness

If the BAM was modified after indexing, the index points to wrong file offsets and region queries return wrong (or zero) reads. Quick check:

if [ input.bam -nt input.bam.bai ]; then
    echo "Index older than BAM; re-indexing"
    samtools index input.bam
fi

Contig-Naming Sanity Check

A leading cause of "my variant calling produced empty VCFs" tickets: querying chrM against a BAM that uses MT (or chr1 vs 1). Always inspect contig conventions before region queries:

samtools view -H input.bam | grep '^@SQ' | head -3
# Compare with reference dict:
samtools dict ref.fa | head -3

UCSC convention uses chr1/chrM; Ensembl/NCBI uses 1/MT. The two are not interchangeable; tools fail with "contig not found" or silently return zero reads.

Common Errors

ErrorCauseSolution
random alignment retrieval only works for indexed BAMMissing indexRun samtools index file.bam
file is not sortedUnsorted BAMSort first with samtools sort
chromosome not foundWrong chromosome nameCheck names with samtools view -H
Region query returns zero reads on a known-populated locusStale BAI / chr vs no-chr mismatchRe-index; verify naming convention
BAI silently truncates reads on contigs >537 MbpPlant / amphibian / amplified genomeUse CSI: samtools index -c file.bam

Related Skills

  • sam-bam-basics - View and convert alignment files
  • alignment-sorting - Sort BAM files (required before indexing)
  • alignment-filtering - Filter by regions using index
  • bam-statistics - Use idxstats for quick counts
  • sequence-io/read-sequences - Index FASTA with SeqIO.index_db()

When not to use it

  • Unsorted alignment files
  • Small files where random access is unnecessary

Prerequisites

samtoolspysam

Limitations

  • Files must be sorted prior to indexing
  • Requires correct tool version identification

How it compares

It automatically chooses the correct index type for large-genome contigs rather than assuming default BAI indexes which fail on large references.

Compared to similar skills

bio-alignment-indexing side by side with the closest alternatives in the catalog.

SkillInstallsUpdatedSafetyDifficulty
bio-alignment-indexing (this skill)02moReviewIntermediate
exploratory-data-analysis152moReviewIntermediate
model-compare77moReviewAdvanced
astropy67moReviewAdvanced

Try saying

Example prompts that trigger this skill in your AI assistant.

You might also like

exploratory-data-analysis

K-Dense-AI

Perform comprehensive exploratory data analysis on scientific data files across 200+ file formats. This skill should be used when analyzing any scientific data file to understand its structure, content, quality, and characteristics. Automatically detects file type and generates detailed markdown reports with format-specific analysis, quality metrics, and downstream analysis recommendations. Covers chemistry, bioinformatics, microscopy, spectroscopy, proteomics, metabolomics, and general scientific data formats.

15114

model-compare

rawwerks

Compare 3D CAD models using boolean operations (IoU, Dice, precision/recall). Use when evaluating generated models against gold references, diffing CAD revisions, or computing similarity metrics for ML training. Triggers on: model diff, compare models, IoU, intersection over union, model similarity, CAD comparison, STEP diff, 3D evaluation, gold reference, generated model, precision recall 3D.

783

astropy

davila7

Comprehensive Python library for astronomy and astrophysics. This skill should be used when working with astronomical data including celestial coordinates, physical units, FITS files, cosmological calculations, time systems, tables, world coordinate systems (WCS), and astronomical data analysis. Use when tasks involve coordinate transformations, unit conversions, FITS file manipulation, cosmological distance calculations, time scale conversions, or astronomical data processing.

682

statistical-analysis

anthropics

Apply statistical methods including descriptive stats, trend analysis, outlier detection, and hypothesis testing. Use when analyzing distributions, testing for significance, detecting anomalies, computing correlations, or interpreting statistical results.

848

datacommons-client

davila7

Work with Data Commons, a platform providing programmatic access to public statistical data from global sources. Use this skill when working with demographic data, economic indicators, health statistics, environmental data, or any public datasets available through Data Commons. Applicable for querying population statistics, GDP figures, unemployment rates, disease prevalence, geographic entity resolution, and exploring relationships between statistical entities.

637

analyzing-market-sentiment

jeremylongshore

Analyze cryptocurrency market sentiment using Fear & Greed Index, news analysis, and market momentum. Use when gauging overall market mood, checking if markets are fearful or greedy, or analyzing sentiment for specific coins. Trigger with phrases like "analyze crypto sentiment", "check market mood", "is the market fearful", "sentiment for Bitcoin", or "Fear and Greed index".

339

Search skills

Search the agent skills registry