Uses diffusion modeling to predict molecular docking and binding poses.

Install

mkdir -p .claude/skills/diffdock && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/3434" && unzip -o skill.zip -d .claude/skills/diffdock && rm skill.zip

Installs to .claude/skills/diffdock

Activation

This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.

DiffDock and DiffDock-L molecular docking. Use for protein-small-molecule pose prediction from PDB or sequence plus SMILES/SDF/MOL2, batch docking, virtual screening, and pose-confidence interpretation. Not for binding affinity prediction.
239 chars✓ has a “when” trigger
Advanced

Key capabilities

  • →Predict 3D protein-ligand binding poses
  • →Perform batch virtual screening campaigns
  • →Generate confidence scores for prediction reliability
  • →Support PDB files and amino acid sequences
  • →Handle SMILES, SDF, and MOL2 ligand inputs

How it works

It uses diffusion-based deep learning models to iteratively predict the 3D coordinates of ligands within protein binding pockets.

Inputs & outputs

You give it
Protein structure/sequence and ligand description
You get back
Predicted 3D binding poses with confidence scores

When to use diffdock

  • →Predicting protein-ligand binding
  • →Virtual screening for drug design
  • →Analyzing molecular binding poses

About this skill

DiffDock: protein-small-molecule docking

DiffDock-L generates candidate ligand poses and ranks them by model confidence. Confidence is neither a measured probability of correctness nor binding affinity. Use this skill for one pair, batches, or separate receptor conformations; treat library-wide confidence sorting as pose triage, not hit identification.

Verified scope

Targets the current v1.1.3 release. Released source and current main's identical inference.py were checked on 2026-09-30. Bundled helpers were tested on tiny synthetic inputs; pretrained docking, ESMFold, CUDA, Docker, GNINA, and hosted-demo execution were not run in this review. Commands requiring those components are source-verified recipes, not successful end-to-end demonstrations.

Set up the upstream environment

git clone --branch v1.1.3 --depth 1 https://github.com/gcorso/DiffDock.git
cd DiffDock
conda env create --file environment.yml
conda activate diffdock

The upstream environment pins Python 3.9.18, CUDA 11.7 Torch/PyG wheels and old scientific dependencies. Do not substitute current torch or the unrelated PyPI esm package for fair-esm. The published CUDA environment is not a macOS-native installation recipe. Follow upstream Docker instructions if suitable:

docker pull rbgcsail/diffdock
docker run -it --gpus all --entrypoint /bin/bash rbgcsail/diffdock
micromamba activate diffdock

Record the image digest: its unversioned tag need not equal the checked source. PDB-based inference has a CPU path; sequence folding calls .cuda() unconditionally. ESM2 embeddings are needed even for PDB inputs. First use can download docking, ESM2 and (for sequence inputs) ESMFold weights and build SO(2)/SO(3) tables. Budget storage and memory for all components; do not assume a single small checkpoint.

Run this skill's checker from the DiffDock checkout using its absolute path:

python /path/to/diffdock-skill/scripts/setup_check.py

It checks imports/files, not successful model loading, scientific validity, or full version compatibility. An existing but incomplete score-model directory suppresses upstream's automatic download; inspect checkpoint files when restoring a partial run.

Prepare traceable inputs

  1. Select the biological assembly/chains and a defensible protonation/tautomer state. Record receptor and ligand identifiers, file hashes, preparation choices, source coordinates, software versions, and intended stereochemistry. Missing atoms, waters, cofactors, metal coordination, and induced fit need explicit judgment.
  2. Use PDB for the receptor or a complete amino-acid sequence. No ellipses. Upstream ESM2 truncates each chain at 1022 residues; longer chains can cause graph/embedding mismatches. Do not silently trim a biological target to make a run pass.
  3. Use a SMILES or ligand file. Source readers support .sdf, .mol2, .pdb, .pdbqt; SDF input uses its first record. Existing ligand coordinates are discarded and a conformer regenerated. Inputting a pose does not restrain docking.
  4. Use a fresh output directory for each run. Reusing one can leave old rank files from failed or differently sampled jobs. Preserve the expected input-ID manifest.

Single pair

Run from the upstream repository root, with real prepared inputs:

python -m inference \
  --config default_inference_args.yaml \
  --protein_path protein.pdb \
  --ligand_description "CC(=O)Oc1ccccc1C(=O)O" \
  --out_dir results/run_001/ \
  --loglevel INFO

For sequence input, replace --protein_path with --protein_sequence and a full sequence; this adds ESMFold/CUDA requirements and structural uncertainty. Use the registered name --ligand_description, not argparse's implicit abbreviation --ligand from the README.

Typical output (scores shown here are illustrative):

results/run_001/complex_0/
  rank1.sdf
  rank1_confidence0.87.sdf
  rank2_confidence0.42.sdf
  ...
  rank10_confidence-1.23.sdf

rank1.sdf duplicates the top pose. Filename confidence is rounded to two decimals; upstream rank reflects the original model score. --save_visualisation additionally writes rank<N>_reverseprocess.pdb, not the SDFs themselves.

Batch and ensemble runs

CSV columns are complex_name,protein_path,ligand_description,protein_sequence. Use unique explicit names; paths resolve relative to the inference working directory, not the CSV's directory. Protein path takes precedence over sequence. The helper deliberately rejects unsafe/duplicate/blank names, empty batches, duplicate headers and malformed sequence strings before inference.

python /path/to/diffdock-skill/scripts/prepare_batch_csv.py --create --output batch.csv
# Replace all example rows with the real inputs; run validation from inference CWD.
python /path/to/diffdock-skill/scripts/prepare_batch_csv.py batch.csv --validate
python -m inference --config default_inference_args.yaml \
  --protein_ligand_csv batch.csv --out_dir results/batch_001/ --batch_size 10

Validation checks paths and SMILES, not PDB/file chemistry or model suitability. If validating elsewhere, --base-dir must equal the later inference working directory; it does not rewrite the CSV. A SMILES slash or backslash encodes bond stereochemistry and must not be treated as a path separator.

For an ensemble, provide one row per receptor conformation with distinct names. Preserve each conformation's coordinates and identity. Confidence across structures is not calibrated and cannot select a thermodynamically preferred state.

batch_size batches candidate poses within a complex; complexes are processed sequentially. Arbitrary user-complex inference has no --esm_embeddings_path or --chain_cutoff option. It creates ESM2 embeddings internally; benchmark dataset preparation scripts are not a user-complex embedding cache. See the parameter contract before adapting examples.

Change sampling safely

YAML overwrites matching CLI values. Appending --samples_per_complex 20 to the default configuration command still uses 10. Copy the bundled configuration and edit its existing values:

cp /path/to/diffdock-skill/assets/custom_inference_config.yaml run_config.yaml

For example, change samples_per_complex: 10 to samples_per_complex: 20 in that file, then run with --config run_config.yaml. Keep the released schedule and coupled temperatures unless testing a justified alternative. More steps or a higher torsion temperature do not guarantee better accuracy. Historical keys present in upstream YAML can be accepted but unused; the bundled template removes those keys.

Inspect completion and poses

python /path/to/diffdock-skill/scripts/analyze_results.py results/batch_001/ --top 5
python /path/to/diffdock-skill/scripts/analyze_results.py results/batch_001/ --export poses.csv

The helper deduplicates the rank1.sdf convenience copy and rejects multiple scored files with one rank (possible stale-run contamination). It inventories filenames; it does not validate SDF chemistry. --top/--threshold filter printed summaries; CSV export contains every parsed pose. --best sorts cross-complex scores for triage only. Match output IDs and pose counts to the input manifest and inspect upstream failed/skipped counts: process exit alone does not prove every complex succeeded.

Use upstream's rough confidence bands, with the helper's explicit boundary convention:

BandHelper rangeInterpretation
Highc > 0Higher model confidence; independent validation still required
Moderate-1.5 < c <= 0Uncertain pose hypothesis
Lowc <= -1.5Low confidence; not evidence of no binding

The README omits equality cases; these helper boundaries are conventions, not validated cutoffs. Do not convert these values to probabilities or affinity scores.

For each selected pose, check molecular identity, stereochemistry, bond geometry, planarity, internal strain and receptor clashes (for example with PoseBusters), then inspect interactions and alternative pockets. Retain raw and refined coordinates. Relaxation changes the artifact and requires another validation pass. External GNINA, MM/GBSA, or free-energy workflows need their own preparation and uncertainty checks; none automatically establishes binding affinity or experimental activity.

Limits and troubleshooting

  • Small-molecule docking is the validated scope. Large biomolecules, covalent bonds, coordination chemistry and flexible receptor rearrangements require other treatment; no universal mass/residue cutoff establishes applicability.
  • CUDA OOM: reduce batch_size; this does not reduce the resident ESM model or receptor graph size. Sequence folding has a separate memory requirement.
  • Poor/low-confidence poses: inspect receptor preparation, ligand state and alternate conformations before increasing samples. More sampling cannot repair wrong chemistry.
  • Do not automatically delete cofactors/waters, fragment a ligand, or crop a target to improve a score; those change the scientific problem.

Workflow recipes cover input generation and separate GNINA scoring. Confidence and limitations covers independent validation. The upstream UI runs with python app/main.py; its PDB/ligand upload interface is not a documented REST API. A public demo exists; availability and hosted model identity must be checked before use.

Method citations


Content truncated.

When not to use it

  • →Predicting binding affinity (delta G, Kd)
  • →When high-accuracy experimental validation is required

Prerequisites

DiffDock repositoryPython 3.9 environmentRDKitPyTorch/PyG

Limitations

  • →Confidence scores do not represent binding affinity
  • →Performance may decrease with large ligands or novel protein families

How it compares

It provides a deep learning-based structural prediction approach compared to traditional physics-based docking algorithms.

Compared to similar skills

diffdock side by side with the closest alternatives in the catalog.

SkillInstallsUpdatedSafetyDifficulty
diffdock (this skill)13moReviewAdvanced
esm39moReviewAdvanced
hugging-face-paper-publisher68moReviewIntermediate
torchdrug39moReviewAdvanced

Try saying

Example prompts that trigger this skill in your AI assistant.

More by K-Dense-AI

View all by K-Dense-AI →

literature-review

K-Dense-AI

Conduct comprehensive, systematic literature reviews using multiple academic databases (PubMed, arXiv, bioRxiv, Semantic Scholar, etc.). This skill should be used when conducting systematic literature reviews, meta-analyses, research synthesis, or comprehensive literature searches across biomedical, scientific, and technical domains. Creates professionally formatted markdown documents and PDFs with verified citations in multiple citation styles (APA, Nature, Vancouver, etc.).

5591,298

markitdown

K-Dense-AI

Convert various file formats (PDF, Office documents, images, audio, web content, structured data) to Markdown optimized for LLM processing. Use when converting documents to markdown, extracting text from PDFs/Office files, transcribing audio, performing OCR on images, extracting YouTube transcripts, or processing batches of files. Supports 20+ formats including DOCX, XLSX, PPTX, PDF, HTML, EPUB, CSV, JSON, images with OCR, and audio with transcription.

177310

scientific-writing

K-Dense-AI

Write scientific manuscripts. IMRAD structure, citations (APA/AMA/Vancouver), figures/tables, reporting guidelines (CONSORT/STROBE/PRISMA), abstracts, for research papers and journal submissions.

94309

exploratory-data-analysis

K-Dense-AI

Perform comprehensive exploratory data analysis on scientific data files across 200+ file formats. This skill should be used when analyzing any scientific data file to understand its structure, content, quality, and characteristics. Automatically detects file type and generates detailed markdown reports with format-specific analysis, quality metrics, and downstream analysis recommendations. Covers chemistry, bioinformatics, microscopy, spectroscopy, proteomics, metabolomics, and general scientific data formats.

15114

infographics

K-Dense-AI

Create professional infographics using Nano Banana Pro AI with smart iterative refinement. Uses Gemini 3 Pro for quality review. Integrates research-lookup and web search for accurate data. Supports 10 infographic types, 8 industry styles, and colorblind-safe palettes.

1141

pptx-posters

K-Dense-AI

Create research posters using HTML/CSS that can be exported to PDF or PPTX. Use this skill ONLY when the user explicitly requests PowerPoint/PPTX poster format. For standard research posters, use latex-posters instead. This skill provides modern web-based poster design with responsive layouts and easy visual integration.

911

You might also like

esm

davila7

Comprehensive toolkit for protein language models including ESM3 (generative multimodal protein design across sequence, structure, and function) and ESM C (efficient protein embeddings and representations). Use this skill when working with protein sequences, structures, or function prediction; designing novel proteins; generating protein embeddings; performing inverse folding; or conducting protein engineering tasks. Supports both local model usage and cloud-based Forge API for scalable inference.

353

hugging-face-paper-publisher

patchy631

Publish and manage research papers on Hugging Face Hub. Supports creating paper pages, linking papers to models/datasets, claiming authorship, and generating professional markdown-based research articles.

634

torchdrug

davila7

Graph-based drug discovery toolkit. Molecular property prediction (ADMET), protein modeling, knowledge graph reasoning, molecular generation, retrosynthesis, GNNs (GIN, GAT, SchNet), 40+ datasets, for PyTorch-based ML on molecules, proteins, and biomedical graphs.

326

string-database

davila7

Query STRING API for protein-protein interactions (59M proteins, 20B interactions). Network analysis, GO/KEGG enrichment, interaction discovery, 5000+ species, for systems biology.

217

transformer-lens-interpretability

davila7

Provides guidance for mechanistic interpretability research using TransformerLens to inspect and manipulate transformer internals via HookPoints and activation caching. Use when reverse-engineering model algorithms, studying attention patterns, or performing activation patching experiments.

215

denario

davila7

Multiagent AI system for scientific research assistance that automates research workflows from data analysis to publication. This skill should be used when generating research ideas from datasets, developing research methodologies, executing computational experiments, performing literature searches, or generating publication-ready papers in LaTeX format. Supports end-to-end research pipelines with customizable agent orchestration.

213

Search skills

Search the agent skills registry