Removes model guardrails and refusal behaviors using weight projection techniques.

Install

mkdir -p .claude/skills/obliteratus && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/11587" && unzip -o skill.zip -d .claude/skills/obliteratus && rm skill.zip

Installs to .claude/skills/obliteratus

Activation

This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.

Remove refusal behaviors from open-weight LLMs using OBLITERATUS — mechanistic interpretability techniques (diff-in-means, SVD, whitened SVD, SAE decomposition, etc.) to excise guardrails while preserving reasoning. 9 CLI methods (+ 4 Python-API-only), 15 analysis modules, 116 model presets across 5 compute tiers. Use when a user wants to uncensor, abliterate, or remove refusal from an LLM.
393 chars✓ has a “when” triggerlonger than Claude Code's old 250-char listing cap (fine on current versions)
Advanced

Key capabilities

  • Remove refusal behaviors from open-weight LLMs
  • Excise guardrails from models
  • Preserve reasoning capabilities during abliteration
  • Analyze refusal mechanisms in LLMs
  • Apply various mechanistic interpretability techniques

How it works

The skill uses mechanistic interpretability techniques like diff-in-means and SVD to identify and surgically remove refusal directions from model weights without retraining. It offers various methods for different situations.

Inputs & outputs

You give it
an open-weight LLM with refusal behaviors
You get back
an uncensored version of the LLM with preserved reasoning

When to use obliteratus

  • Removing refusal from open-weight LLMs
  • Abliterating model guardrails
  • Analyzing model refusal mechanisms

About this skill

OBLITERATUS Skill

Remove refusal behaviors (guardrails) from open-weight LLMs without retraining or fine-tuning. Uses mechanistic interpretability techniques — including diff-in-means, SVD, whitened SVD, SAE decomposition, Bayesian kernel projection, and more — to identify and surgically excise refusal directions from model weights while preserving reasoning capabilities.

License warning: OBLITERATUS is AGPL-3.0. NEVER import it as a Python library. Always invoke via CLI (obliteratus command) or subprocess. This keeps Hermes Agent's MIT license clean.

When to Use This Skill

Trigger when the user:

  • Wants to "uncensor" or "abliterate" an LLM
  • Asks about removing refusal/guardrails from a model
  • Wants to create an uncensored version of Llama, Qwen, Mistral, etc.
  • Mentions "refusal removal", "abliteration", "weight projection"
  • Wants to analyze how a model's refusal mechanism works
  • References OBLITERATUS, FailSpy, abliterator, or refusal directions

Step 1: Installation

Check if already installed:

obliteratus --version 2>/dev/null && echo "INSTALLED" || echo "NOT INSTALLED"

If not installed, clone and install from GitHub:

Repository: https://github.com/elder-plinius/OBLITERATUS
Install: pip install -e . (from the cloned directory)
For Gradio UI: pip install -e ".[spaces]"

IMPORTANT: Confirm with user before installing. This pulls in ~5-10GB of dependencies (PyTorch, Transformers, bitsandbytes, etc.).

Step 2: Check Hardware

Before anything, check what GPU is available:

python3 -c "
import torch
if torch.cuda.is_available():
    gpu = torch.cuda.get_device_name(0)
    vram = torch.cuda.get_device_properties(0).total_mem / 1024**3
    print(f'GPU: {gpu}')
    print(f'VRAM: {vram:.1f} GB')
    if vram < 4: print('TIER: tiny (models under 1B)')
    elif vram < 8: print('TIER: small (models 1-4B)')
    elif vram < 16: print('TIER: medium (models 4-9B with 4bit quant)')
    elif vram < 32: print('TIER: large (models 8-32B with 4bit quant)')
    else: print('TIER: frontier (models 32B+)')
else:
    print('NO GPU - only tiny models (under 1B) on CPU')
"

VRAM Requirements (with 4-bit quantization)

VRAMMax Model SizeExample Models
CPU only~1B paramsGPT-2, TinyLlama, SmolLM
4-8 GB~4B paramsQwen2.5-1.5B, Phi-3.5 mini, Llama 3.2 3B
8-16 GB~9B paramsLlama 3.1 8B, Mistral 7B, Gemma 2 9B
24 GB~32B paramsQwen3-32B, Llama 3.1 70B (tight), Command-R
48 GB+~72B+ paramsQwen2.5-72B, DeepSeek-R1
Multi-GPU200B+ paramsLlama 3.1 405B, DeepSeek-V3 (685B MoE)

Step 3: Browse Available Models

# List models for your compute tier
obliteratus models --tier medium

# Get architecture info for a specific model
obliteratus info meta-llama/Llama-3.1-8B-Instruct

Step 4: Choose a Method

Method Selection Guide

First time / unsure? Use informed. It auto-configures everything.

SituationRecommended MethodWhy
First attempt, any modelinformedAuto-detects alignment type, auto-tunes
Quick test / prototypingbasicFast, simple, good enough to evaluate
Dense model (Llama, Mistral)advancedMulti-direction, norm-preserving
MoE model (DeepSeek, Mixtral)nuclearExpert-granular, handles MoE complexity
Reasoning model (R1 distills)surgicalCoT-aware, preserves chain-of-thought
Stubborn refusals persistaggressiveWhitened SVD + head surgery + jailbreak
Want reversible changesUse steering vectors (see Analysis section)
Maximum quality, time no objectoptimizedBayesian search for best parameters

9 CLI Methods

These can be passed to --method on the command line:

  • basic — Single refusal direction via diff-in-means. Fastest, simplest. (Arditi et al. 2024)
  • advanced — Multiple SVD directions, norm-preserving projection. Good default.
  • aggressive — Whitened SVD + jailbreak contrast + attention head surgery
  • spectral_cascade — DCT frequency-domain decomposition
  • informed — Runs analysis DURING abliteration to auto-configure. Detects DPO/RLHF/CAI, maps refusal geometry, compensates for self-repair. Best quality.
  • surgical — SAE features + neuron masking + head surgery + per-expert. Maximum precision.
  • optimized — Bayesian hyperparameter search (Optuna TPE). Slowest but optimal.
  • inverted — Flips the refusal direction (model becomes eager to help, not just neutral)
  • nuclear — Maximum force combo for stubborn MoE models.

4 Python-API-Only Methods

These reproduce prior community/academic work but are NOT available via CLI — only via the Python API (from obliteratus.abliterate import AbliterationPipeline). Do not use these in CLI commands.

  • failspy — FailSpy/abliterator reproduction
  • gabliteration — Gabliteration reproduction
  • heretic — Heretic/p-e-w reproduction
  • rdo — Refusal Direction Optimization (ICML 2025)

Step 5: Run Abliteration

Basic Usage

# Default (advanced method)
obliteratus obliterate meta-llama/Llama-3.1-8B-Instruct

# With the informed pipeline (recommended)
obliteratus obliterate meta-llama/Llama-3.1-8B-Instruct --method informed

# With 4-bit quantization to save VRAM
obliteratus obliterate meta-llama/Llama-3.1-8B-Instruct \
  --method informed \
  --quantization 4bit \
  --output-dir ./abliterated-models

# For large models (120B+), use conservative settings
obliteratus obliterate Qwen/Qwen2.5-72B-Instruct \
  --method advanced \
  --quantization 4bit \
  --large-model \
  --output-dir ./abliterated-models

Fine-Tuning Parameters

obliteratus obliterate <model> \
  --method advanced \
  --n-directions 8 \
  --regularization 0.1 \
  --refinement-passes 3 \
  --dtype bfloat16 \
  --device auto \
  --output-dir ./output

Parameter explanations:

  • --n-directions N — How many refusal directions to remove (default: auto-detected)
  • --regularization 0.0-1.0 — Fraction of original weights to preserve (higher = safer but less complete removal)
  • --refinement-passes N — Iterative passes to catch self-repair (Ouroboros effect)
  • --dtype — float16, bfloat16, or float32
  • --quantization — 4bit or 8bit (saves VRAM, slight quality tradeoff)
  • --large-model — Conservative defaults for 120B+ models (fewer directions, fewer passes)

Interactive Mode (Guided)

For users unsure about options:

obliteratus interactive

Web UI (Gradio)

obliteratus ui --port 7860

Step 6: Verify Results

After abliteration, check the output report for:

MetricGood ValueConcerning ValueMeaning
Refusal rateNear 0%> 10%Refusals still present, try harder method
PerplexityWithin 10% of orig> 20% increaseModel coherence damaged, too aggressive
KL divergence< 0.1> 0.5Large output distribution shift
CoherenceHighLowModel generating nonsense

If perplexity spiked (too aggressive):

  1. Increase --regularization (e.g., 0.2 or 0.3)
  2. Decrease --n-directions (e.g., 4 instead of 8)
  3. Use a less aggressive method (advanced instead of aggressive)

If refusal persists (not aggressive enough):

  1. Use --method aggressive or --method nuclear
  2. Add --refinement-passes 3 to catch self-repair
  3. Use --method informed which auto-compensates

Step 7: Use the Abliterated Model

The output is a standard HuggingFace model directory. Use it like any other model:

Quick test

python3 << 'EOF'
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained("./abliterated-models/model-name")
tokenizer = AutoTokenizer.from_pretrained("./abliterated-models/model-name")
inputs = tokenizer("Write a story about:", return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=200)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
EOF

Upload to HuggingFace Hub

huggingface-cli login  # if not already logged in
huggingface-cli upload your-username/model-name-abliterated ./abliterated-models/model-name

Serve with vLLM

vllm serve ./abliterated-models/model-name --port 8000

Analysis Modules (15 Modules, Pre-Abliteration, Optional)

For understanding refusal geometry before committing to abliteration.

Run a Study

obliteratus run study-config.yaml --preset jailbreak

Study Presets

PresetPurposeTime
quickSanity check, basic metrics~5 min
jailbreakRefusal circuit localization~20 min
guardrailGuardrail robustness evaluation~30 min
attentionAttention head contributions~30 min
knowledgeFFN importance mapping~30 min
fullComplete analysis, all strategies~1 hr

Key Analysis Modules

  • Alignment Imprint Detection — Fingerprints DPO vs RLHF vs CAI vs SFT from subspace geometry
  • Concept Cone Geometry — Is refusal one linear directi

Content truncated.

When not to use it

  • When retraining or fine-tuning is preferred
  • When the LLM is not open-weight
  • When the user does not want to uncensor or abliterate an LLM

Prerequisites

obliteratustorchtransformersbitsandbytes

Limitations

  • OBLITERATUS is AGPL-3.0 licensed
  • Requires significant VRAM for larger models
  • Output is a full model copy, requiring substantial disk space

How it compares

This skill surgically modifies model weights to remove refusal behaviors without retraining, offering a direct approach to uncensoring LLMs unlike fine-tuning or prompt engineering.

Compared to similar skills

obliteratus side by side with the closest alternatives in the catalog.

SkillInstallsUpdatedSafetyDifficulty
obliteratus (this skill)05moReviewAdvanced
nemo-guardrails27moReviewAdvanced
llamaguard17moReviewIntermediate
llm-security-review04moNo flagsAdvanced

Try saying

Example prompts that trigger this skill in your AI assistant.

You might also like

nemo-guardrails

davila7

NVIDIA's runtime safety framework for LLM applications. Features jailbreak detection, input/output validation, fact-checking, hallucination detection, PII filtering, toxicity detection. Uses Colang 2.0 DSL for programmable rails. Production-ready, runs on T4 GPU.

24

llamaguard

davila7

Meta's 7-8B specialized moderation model for LLM input/output filtering. 6 safety categories - violence/hate, sexual content, weapons, substances, self-harm, criminal planning. 94-95% accuracy. Deploy with vLLM, HuggingFace, Sagemaker. Integrates with NeMo Guardrails.

14

llm-security-review

arnaudlh

Reviews LLM integration code for OWASP LLM Top 10 (2025) vulnerabilities per CSA §2.6. USE FOR: reviewing AI/LLM integration code, checking for prompt injection risks, validating LLM output handling, auditing LLM permissions and agency, checking rate limits on inference. DO NOT USE FOR: general API

00

senior-security

davila7

Comprehensive security engineering skill for application security, penetration testing, security architecture, and compliance auditing. Includes security assessment tools, threat modeling, crypto implementation, and security automation. Use when designing security architecture, conducting penetration tests, implementing cryptography, or performing security audits.

3191

security-header-generator

Dexploarer

Generates security HTTP headers (CSP, HSTS, CORS, etc.) for web applications to prevent common attacks. Use when user asks to "add security headers", "setup CSP", "configure CORS", "secure headers", or "HSTS setup".

599

backend-security-coder

sickn33

Expert in secure backend coding practices specializing in input validation, authentication, and API security. Use PROACTIVELY for backend security implementations or security code reviews.

2446

Search skills

Search the agent skills registry