BI

bio-phasing-imputation-haplotype-phasing

Phases genotype data into haplotypes to prepare for genetic analysis and population studies.

Install

mkdir -p .claude/skills/bio-phasing-imputation-haplotype-phasing && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/16081" && unzip -o skill.zip -d .claude/skills/bio-phasing-imputation-haplotype-phasing && rm skill.zip

Installs to .claude/skills/bio-phasing-imputation-haplotype-phasing

Activation

This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.

<!-- # COPYRIGHT NOTICE # This file is part of the "Universal Biomedical Skills" project. # Copyright (c) 2026 MD BABU MIA, PhD <[email protected]> # All Rights Reserved. # # This code is proprietary and confidential. # Unauthorized copying of this file, via any medium is stri
280 chars · catalog descriptionno explicit “when” triggerlonger than Claude Code's old 250-char listing cap (fine on current versions)
Advanced

Key capabilities

  • →Phase genotypes into haplotypes
  • →Prepare VCF files for imputation
  • →Use Beagle for phasing
  • →Use SHAPEIT5 for phasing large datasets
  • →Check phasing results for phased vs unphased genotypes

How it works

It uses CLI tools like Beagle or SHAPEIT5 to resolve allele inheritance, processing VCF files to create phased haplotype data, optionally using genetic maps or reference panels for accuracy.

Inputs & outputs

You give it
VCF file (e.g., input.vcf.gz), optional genetic map, optional reference panel
You get back
phased VCF file (e.g., phased.vcf.gz) with phased genotypes

When to use bio-phasing-imputation-haplotype-phasing

  • →Preparing VCF files
  • →Haplotype phasing
  • →Population genetic analysis

About this skill

Version Compatibility

Reference examples tested with: SHAPEIT5 5.1.1, Eagle 2.4.1, Beagle 5.4 (22Jul22), bcftools 1.19+.

Before using code patterns, verify installed versions match. If versions differ:

  • CLI: <tool> --version then <tool> --help to confirm flags

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

SHAPEIT4 to SHAPEIT5 changed the CLI substantially: SHAPEIT5 is a SUITE of binaries (phase_common, phase_rare, ligate, switch), not a single shapeit command, and phase_common is the engine formerly known as SHAPEIT4. The genetic map and the reference panel must match the data's genome build (GRCh37 vs GRCh38); a build-mismatched map silently degrades phasing. PBWT and Ne defaults have drifted between betas; confirm against the installed --help.

Statistical Haplotype Phasing -- Inferring Phase From Population LD

"Resolve which alleles sit together on each chromosome" -> Estimate haplotype phase from population linkage disequilibrium via the Li-Stephens HMM - because phase is INFERRED statistically from how haplotypes are shared across a population, not read off the genotype, so a switch error is a model uncertainty (the rate, not zero, is the deliverable), not a typo.

  • CLI: phase_common --input target.bcf --filter-maf 0.001 --map chr20.b38.gmap.gz --region chr20 --output scaffold.bcf then ligate then phase_rare (SHAPEIT5), or Eagle2/Beagle for common-variant phasing

Scope: population/statistical phasing of array or sequence genotypes for imputation input, compound-het/ASE/HLA, and population genetics. Read-backed / molecular single-sample phasing (long reads, Hi-C, 10x linked reads) is a PHYSICALLY DIFFERENT signal -> long-read-sequencing/haplotype-phasing (the two are easily conflated; do not run SHAPEIT on long-read evidence or trust statistical phase for a private clinical variant). Panel choice -> reference-panels. Imputation against a panel -> genotype-imputation. The input VCF and biallelic normalization -> variant-calling/variant-normalization. End-to-end orchestration -> workflows/gwas-pipeline.

The Single Most Important Modern Insight -- A Phased Haplotype Is a Statistical Estimate, and Its Error Concentrates Exactly Where the Biology of Interest Lives

Statistical phasing reconstructs which alleles are on the same chromosome by borrowing LD across many individuals or a reference panel (Delaneau 2019 Nat Commun 10:5436). That works beautifully for common variants in LD with their neighbors and fails, by construction, for rare variants - which are young, carried by few people, and in LD with almost nothing. Three facts drive every decision:

  1. The genome-wide switch-error rate lies, because it is dominated by easy common sites. A headline "switch error rate 0.3%" is averaged over millions of common heterozygous sites and says nothing about the singleton or doubleton that is most likely to be the compound-het, the de-novo, or the pathogenic allele of interest - those are phased at MAC-dependent accuracy an order of magnitude worse, and a true singleton is essentially a coin flip without special machinery (Hofmeister 2023 Nat Genet 55:1243). Report accuracy stratified by minor allele count, never as one number.
  2. The deliverable is a switch-error rate against an independent truth set, not the tool name. "We used SHAPEIT" is not a switch-error rate. A switch error changes which haplotype an allele sits on without changing any genotype, so it is invisible to every per-site genotype QC; for any phase-dependent claim, measure the rate against a trio (Mendelian truth via switch --pedigree) or read-backed truth.
  3. The modern arc is the scaffold design, and rare-variant phasing needs biobank scale to work at all. SHAPEIT5 phases common variants into a fixed, near-perfect scaffold, then places each rare allele onto it by PBWT/IBD haplotype matching - which depends on finding a long shared haplotype, itself a function of cohort size. This is why rare-variant phasing in a small cohort cannot be trusted for a cis/trans call without orthogonal (trio or read-backed) evidence.

Tool Taxonomy

ToolCitationMechanism / roleWhen
SHAPEIT5Hofmeister 2023 Nat Genet 55:1243suite (phase_common/phase_rare/ligate/switch); scaffold design for rare/singleton phasing; PBWTbiobank-scale WGS/WES; rare-variant phasing
SHAPEIT4 (= phase_common engine)Delaneau 2019 Nat Commun 10:5436sub-linear common-variant phasing; integrates panels, scaffolds, read-backed phasecommon-variant phasing / pre-phasing; legacy
Eagle2Loh 2016 Nat Genet 48:1443HMM + PBWT-derived HapHedge; reference-based (--vcfRef) and within-cohortarray data; the classic imputation-server phaser
Beagle 5.xBrowning 2021 Am J Hum Genet 108:1880Java; does BOTH phasing (gt=, no ref=) and imputation; two-stage for sequenceone tool for phase and impute; no compile
Trio / pedigree phasing(Mendelian transmission)deterministic phase where the trio is informativegold standard; validating other phasers via switch
WhatsHap (boundary)Patterson 2015 J Comput Biol 22:498read-backed phasing (weighted MEC) from aligned reads-> long-read-sequencing/haplotype-phasing; can seed SHAPEIT as a scaffold

Decision Tree by Scenario

ScenarioRecommendedWhy
Array data, small-to-modest cohort, have a panelEagle2 --vcfRef or phase_common --referencea panel models LD better than a few thousand samples
Array data, large cohort, no panelEagle2 or phase_common within-cohortLD is modeled from the cohort; accuracy rises with N
WGS/WES, biobank scale, need rare variants phasedSHAPEIT5: phase_common -> ligate -> phase_rarethe scaffold design is the only route to accurate rare-variant phase
Pre-phasing as imputation inputEagle2 or Beagle 5small switch errors largely wash out in imputation -> genotype-imputation
One tool for phase and impute, no compileBeagle 5.x (gt= to phase, add ref= to impute)pragmatic single tool
Trio/pedigree availabletrio/pedigree phasing; use switch to benchmarkdeterministic where informative; the truth ruler
Long reads on the same sample-> long-read-sequencing/haplotype-phasing (then seed SHAPEIT as a scaffold)read-backed phase is local and deterministic; combine, do not replace
Common-variant phasing only, modest dataSHAPEIT4 or Beaglerare-variant machinery is unnecessary overhead

The Common-Scaffold-Then-Rare Design (SHAPEIT5)

Rare variants carry too little LD to phase in a joint model, and a joint HMM over millions of rare sites does not scale, so SHAPEIT5 splits the problem. Use the full pipeline when N > ~2,000; below that, phase_common alone suffices (too few rare-allele carriers for the rare step to add value).

  1. phase_common phases the common variants (e.g. --filter-maf 0.001) into accurate haplotypes - the scaffold. Run per chunk for large chromosomes, with OVERLAPPING regions.
  2. ligate stitches the per-chunk common scaffolds into one chromosome; chunks must overlap so ligate can resolve phase across the seam (a non-overlapping seam is a guaranteed switch).
  3. phase_rare takes the FULL genotypes plus the fixed scaffold and places each rare allele onto the already-phased common haplotypes by IBD matching. Do not filter rare variants out of the phase_rare input - placing them is the whole point.

Switch Error vs Flip vs Hamming -- the Metrics

A single rate hides the failure mode. Report more than one, and look at the distribution of switch positions.

MetricWhat it countsInflates on
Switch error rate (SER)fraction of consecutive het-site pairs whose phase relationship is wrongmany small local errors; the standard headline
Flip erroran isolated het phased wrong then immediately corrected (two switches one site apart)noisy single sites; double-counts in raw SER
Hamming errorfraction of het sites on the wrong haplotype under the best global alignmenta few LARGE block swaps - high Hamming, low switch count
Long switch / block flipa sustained segment on the wrong haplotypepoor long-range LD; ruinous for cis/trans yet only 2 switches

SER and Hamming measure different sins: many tiny flips give high SER but modest Hamming; one half-chromosome block swap gives catastrophic Hamming but only two switches. Het density matters too - SER is per-het-pair, so sparse het sites mean the same SER spans more bp. Typical magnitudes (order-of-magnitude, dataset-specific): Eagle2 + HRC reference, European array 1.36%; Eagle2 within-cohort N5,000 1.5%; within-cohort N150,000 (UK Biobank) ~0.27-0.35%; SHAPEIT5 for a variant in ~1 of 100,000 < ~5%. The pattern: common-variant phasing in a big cohort is sub-1%; rare-variant phasing is single-digit-percent at best and worsens steeply as MAC approaches 1.

Reference-Based vs Within-Cohort

Reference-based phasing wins when the cohort is small (a few thousand samples cannot model LD as well as a 32k-100k+ haplotype panel); phase against the biggest ancestry-matched panel available (Eagle2 --vcfRef). Within-cohort phasing wins when the cohort is large and ancestry-matched to itself, because accuracy rises monotonically with N; by UK-Biobank scale within-cohort is more accurate than any external panel. The crossover is in the tens of thousands. Ancestry match dominates either way - a mismatched panel phases worse than a smaller matched one or within-cohort -> reference-panels.

Per-Method Failure Modes

Genome-wide SER trusted for a rare-variant call

Trigger: quoting one switch-error rate and treating all haplotypes as equally trustworthy. Mechanism: SER is dom


Content truncated.

When not to use it

  • →When the task requires genotype imputation before phasing
  • →When the task requires reference data not available in reference panels

Limitations

  • →Beagle memory requirements can be 64+ GB for 100,000 samples
  • →SHAPEIT5 memory requirements can be 32 GB for 100,000 samples
  • →Genetic maps are required for improved accuracy

How it compares

This skill provides a specialized workflow for haplotype phasing using established bioinformatics tools, ensuring accurate resolution of allele inheritance for downstream genetic analyses.

Compared to similar skills

bio-phasing-imputation-haplotype-phasing side by side with the closest alternatives in the catalog.

SkillInstallsUpdatedSafetyDifficulty
bio-phasing-imputation-haplotype-phasing (this skill)06moReviewAdvanced
jupyter-notebook307moReviewIntermediate
juicebox-core-workflow-b12moReviewAdvanced
obspy-data-api18moNo flagsIntermediate

Try saying

Example prompts that trigger this skill in your AI assistant.

You might also like

jupyter-notebook

davila7

Use when the user asks to create, scaffold, or edit Jupyter notebooks (`.ipynb`) for experiments, explorations, or tutorials; prefer the bundled templates and run the helper script `new_notebook.py` to generate a clean starting notebook.

30158

juicebox-core-workflow-b

jeremylongshore

Implement Juicebox candidate enrichment workflow. Use when enriching profile data, gathering additional candidate details, or building comprehensive candidate profiles. Trigger with phrases like "juicebox enrich profile", "juicebox candidate details", "enrich candidate data", "juicebox profile enrichment".

13

obspy-data-api

benchflow-ai

An overview of the core data API of ObsPy, a Python framework for processing seismological data. It is useful for parsing common seismological file formats, or manipulating custom data into standard objects for downstream use cases such as ObsPy's signal processing routines or SeisBench's modeling API.

12

perplexity-core-workflow-b

jeremylongshore

Execute Perplexity secondary workflow: Core Workflow B. Use when implementing secondary use case, or complementing primary workflow. Trigger with phrases like "perplexity secondary workflow", "secondary task with perplexity".

11

source-coding

parcadei

Problem-solving strategies for source coding in information theory

11

biopython

davila7

Primary Python toolkit for molecular biology. Preferred for Python-based PubMed/NCBI queries (Bio.Entrez), sequence manipulation, file parsing (FASTA, GenBank, FASTQ, PDB), advanced BLAST workflows, structures, phylogenetics. For quick BLAST, use gget. For direct REST API, use pubmed-database.

10

Search skills

Search the agent skills registry