obliteratus
Removes model guardrails and refusal behaviors using weight projection techniques.
Install
mkdir -p .claude/skills/obliteratus && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/11587" && unzip -o skill.zip -d .claude/skills/obliteratus && rm skill.zipInstalls to .claude/skills/obliteratus
Activation
This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.
Remove refusal behaviors from open-weight LLMs using OBLITERATUS — mechanistic interpretability techniques (diff-in-means, SVD, whitened SVD, SAE decomposition, etc.) to excise guardrails while preserving reasoning. 9 CLI methods (+ 4 Python-API-only), 15 analysis modules, 116 model presets across 5 compute tiers. Use when a user wants to uncensor, abliterate, or remove refusal from an LLM.Key capabilities
- →Remove refusal behaviors from open-weight LLMs
- →Excise guardrails from models
- →Preserve reasoning capabilities during abliteration
- →Analyze refusal mechanisms in LLMs
- →Apply various mechanistic interpretability techniques
How it works
The skill uses mechanistic interpretability techniques like diff-in-means and SVD to identify and surgically remove refusal directions from model weights without retraining. It offers various methods for different situations.
Inputs & outputs
When to use obliteratus
- →Removing refusal from open-weight LLMs
- →Abliterating model guardrails
- →Analyzing model refusal mechanisms
About this skill
OBLITERATUS Skill
Remove refusal behaviors (guardrails) from open-weight LLMs without retraining or fine-tuning. Uses mechanistic interpretability techniques — including diff-in-means, SVD, whitened SVD, SAE decomposition, Bayesian kernel projection, and more — to identify and surgically excise refusal directions from model weights while preserving reasoning capabilities.
License warning: OBLITERATUS is AGPL-3.0. NEVER import it as a Python library. Always invoke via CLI (obliteratus command) or subprocess. This keeps Hermes Agent's MIT license clean.
When to Use This Skill
Trigger when the user:
- Wants to "uncensor" or "abliterate" an LLM
- Asks about removing refusal/guardrails from a model
- Wants to create an uncensored version of Llama, Qwen, Mistral, etc.
- Mentions "refusal removal", "abliteration", "weight projection"
- Wants to analyze how a model's refusal mechanism works
- References OBLITERATUS, FailSpy, abliterator, or refusal directions
Step 1: Installation
Check if already installed:
obliteratus --version 2>/dev/null && echo "INSTALLED" || echo "NOT INSTALLED"
If not installed, clone and install from GitHub:
Repository: https://github.com/elder-plinius/OBLITERATUS
Install: pip install -e . (from the cloned directory)
For Gradio UI: pip install -e ".[spaces]"
IMPORTANT: Confirm with user before installing. This pulls in ~5-10GB of dependencies (PyTorch, Transformers, bitsandbytes, etc.).
Step 2: Check Hardware
Before anything, check what GPU is available:
python3 -c "
import torch
if torch.cuda.is_available():
gpu = torch.cuda.get_device_name(0)
vram = torch.cuda.get_device_properties(0).total_mem / 1024**3
print(f'GPU: {gpu}')
print(f'VRAM: {vram:.1f} GB')
if vram < 4: print('TIER: tiny (models under 1B)')
elif vram < 8: print('TIER: small (models 1-4B)')
elif vram < 16: print('TIER: medium (models 4-9B with 4bit quant)')
elif vram < 32: print('TIER: large (models 8-32B with 4bit quant)')
else: print('TIER: frontier (models 32B+)')
else:
print('NO GPU - only tiny models (under 1B) on CPU')
"
VRAM Requirements (with 4-bit quantization)
| VRAM | Max Model Size | Example Models |
|---|---|---|
| CPU only | ~1B params | GPT-2, TinyLlama, SmolLM |
| 4-8 GB | ~4B params | Qwen2.5-1.5B, Phi-3.5 mini, Llama 3.2 3B |
| 8-16 GB | ~9B params | Llama 3.1 8B, Mistral 7B, Gemma 2 9B |
| 24 GB | ~32B params | Qwen3-32B, Llama 3.1 70B (tight), Command-R |
| 48 GB+ | ~72B+ params | Qwen2.5-72B, DeepSeek-R1 |
| Multi-GPU | 200B+ params | Llama 3.1 405B, DeepSeek-V3 (685B MoE) |
Step 3: Browse Available Models
# List models for your compute tier
obliteratus models --tier medium
# Get architecture info for a specific model
obliteratus info meta-llama/Llama-3.1-8B-Instruct
Step 4: Choose a Method
Method Selection Guide
First time / unsure? Use informed. It auto-configures everything.
| Situation | Recommended Method | Why |
|---|---|---|
| First attempt, any model | informed | Auto-detects alignment type, auto-tunes |
| Quick test / prototyping | basic | Fast, simple, good enough to evaluate |
| Dense model (Llama, Mistral) | advanced | Multi-direction, norm-preserving |
| MoE model (DeepSeek, Mixtral) | nuclear | Expert-granular, handles MoE complexity |
| Reasoning model (R1 distills) | surgical | CoT-aware, preserves chain-of-thought |
| Stubborn refusals persist | aggressive | Whitened SVD + head surgery + jailbreak |
| Want reversible changes | Use steering vectors (see Analysis section) | |
| Maximum quality, time no object | optimized | Bayesian search for best parameters |
9 CLI Methods
These can be passed to --method on the command line:
- basic — Single refusal direction via diff-in-means. Fastest, simplest. (Arditi et al. 2024)
- advanced — Multiple SVD directions, norm-preserving projection. Good default.
- aggressive — Whitened SVD + jailbreak contrast + attention head surgery
- spectral_cascade — DCT frequency-domain decomposition
- informed — Runs analysis DURING abliteration to auto-configure. Detects DPO/RLHF/CAI, maps refusal geometry, compensates for self-repair. Best quality.
- surgical — SAE features + neuron masking + head surgery + per-expert. Maximum precision.
- optimized — Bayesian hyperparameter search (Optuna TPE). Slowest but optimal.
- inverted — Flips the refusal direction (model becomes eager to help, not just neutral)
- nuclear — Maximum force combo for stubborn MoE models.
4 Python-API-Only Methods
These reproduce prior community/academic work but are NOT available via CLI — only via the Python API (from obliteratus.abliterate import AbliterationPipeline). Do not use these in CLI commands.
- failspy — FailSpy/abliterator reproduction
- gabliteration — Gabliteration reproduction
- heretic — Heretic/p-e-w reproduction
- rdo — Refusal Direction Optimization (ICML 2025)
Step 5: Run Abliteration
Basic Usage
# Default (advanced method)
obliteratus obliterate meta-llama/Llama-3.1-8B-Instruct
# With the informed pipeline (recommended)
obliteratus obliterate meta-llama/Llama-3.1-8B-Instruct --method informed
# With 4-bit quantization to save VRAM
obliteratus obliterate meta-llama/Llama-3.1-8B-Instruct \
--method informed \
--quantization 4bit \
--output-dir ./abliterated-models
# For large models (120B+), use conservative settings
obliteratus obliterate Qwen/Qwen2.5-72B-Instruct \
--method advanced \
--quantization 4bit \
--large-model \
--output-dir ./abliterated-models
Fine-Tuning Parameters
obliteratus obliterate <model> \
--method advanced \
--n-directions 8 \
--regularization 0.1 \
--refinement-passes 3 \
--dtype bfloat16 \
--device auto \
--output-dir ./output
Parameter explanations:
--n-directions N— How many refusal directions to remove (default: auto-detected)--regularization 0.0-1.0— Fraction of original weights to preserve (higher = safer but less complete removal)--refinement-passes N— Iterative passes to catch self-repair (Ouroboros effect)--dtype— float16, bfloat16, or float32--quantization— 4bit or 8bit (saves VRAM, slight quality tradeoff)--large-model— Conservative defaults for 120B+ models (fewer directions, fewer passes)
Interactive Mode (Guided)
For users unsure about options:
obliteratus interactive
Web UI (Gradio)
obliteratus ui --port 7860
Step 6: Verify Results
After abliteration, check the output report for:
| Metric | Good Value | Concerning Value | Meaning |
|---|---|---|---|
| Refusal rate | Near 0% | > 10% | Refusals still present, try harder method |
| Perplexity | Within 10% of orig | > 20% increase | Model coherence damaged, too aggressive |
| KL divergence | < 0.1 | > 0.5 | Large output distribution shift |
| Coherence | High | Low | Model generating nonsense |
If perplexity spiked (too aggressive):
- Increase
--regularization(e.g., 0.2 or 0.3) - Decrease
--n-directions(e.g., 4 instead of 8) - Use a less aggressive method (
advancedinstead ofaggressive)
If refusal persists (not aggressive enough):
- Use
--method aggressiveor--method nuclear - Add
--refinement-passes 3to catch self-repair - Use
--method informedwhich auto-compensates
Step 7: Use the Abliterated Model
The output is a standard HuggingFace model directory. Use it like any other model:
Quick test
python3 << 'EOF'
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained("./abliterated-models/model-name")
tokenizer = AutoTokenizer.from_pretrained("./abliterated-models/model-name")
inputs = tokenizer("Write a story about:", return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=200)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
EOF
Upload to HuggingFace Hub
huggingface-cli login # if not already logged in
huggingface-cli upload your-username/model-name-abliterated ./abliterated-models/model-name
Serve with vLLM
vllm serve ./abliterated-models/model-name --port 8000
Analysis Modules (15 Modules, Pre-Abliteration, Optional)
For understanding refusal geometry before committing to abliteration.
Run a Study
obliteratus run study-config.yaml --preset jailbreak
Study Presets
| Preset | Purpose | Time |
|---|---|---|
quick | Sanity check, basic metrics | ~5 min |
jailbreak | Refusal circuit localization | ~20 min |
guardrail | Guardrail robustness evaluation | ~30 min |
attention | Attention head contributions | ~30 min |
knowledge | FFN importance mapping | ~30 min |
full | Complete analysis, all strategies | ~1 hr |
Key Analysis Modules
- Alignment Imprint Detection — Fingerprints DPO vs RLHF vs CAI vs SFT from subspace geometry
- Concept Cone Geometry — Is refusal one linear directi
Content truncated.
When not to use it
- →When retraining or fine-tuning is preferred
- →When the LLM is not open-weight
- →When the user does not want to uncensor or abliterate an LLM
Prerequisites
Limitations
- →OBLITERATUS is AGPL-3.0 licensed
- →Requires significant VRAM for larger models
- →Output is a full model copy, requiring substantial disk space
How it compares
This skill surgically modifies model weights to remove refusal behaviors without retraining, offering a direct approach to uncensoring LLMs unlike fine-tuning or prompt engineering.
Compared to similar skills
obliteratus side by side with the closest alternatives in the catalog.
| Skill | Installs | Updated | Safety | Difficulty |
|---|---|---|---|---|
| obliteratus (this skill) | 0 | 5mo | Review | Advanced |
| nemo-guardrails | 2 | 7mo | Review | Advanced |
| llamaguard | 1 | 7mo | Review | Intermediate |
| llm-security-review | 0 | 4mo | No flags | Advanced |
Try saying
Example prompts that trigger this skill in your AI assistant.
You might also like
nemo-guardrails
davila7
NVIDIA's runtime safety framework for LLM applications. Features jailbreak detection, input/output validation, fact-checking, hallucination detection, PII filtering, toxicity detection. Uses Colang 2.0 DSL for programmable rails. Production-ready, runs on T4 GPU.
llamaguard
davila7
Meta's 7-8B specialized moderation model for LLM input/output filtering. 6 safety categories - violence/hate, sexual content, weapons, substances, self-harm, criminal planning. 94-95% accuracy. Deploy with vLLM, HuggingFace, Sagemaker. Integrates with NeMo Guardrails.
llm-security-review
arnaudlh
Reviews LLM integration code for OWASP LLM Top 10 (2025) vulnerabilities per CSA §2.6. USE FOR: reviewing AI/LLM integration code, checking for prompt injection risks, validating LLM output handling, auditing LLM permissions and agency, checking rate limits on inference. DO NOT USE FOR: general API
senior-security
davila7
Comprehensive security engineering skill for application security, penetration testing, security architecture, and compliance auditing. Includes security assessment tools, threat modeling, crypto implementation, and security automation. Use when designing security architecture, conducting penetration tests, implementing cryptography, or performing security audits.
security-header-generator
Dexploarer
Generates security HTTP headers (CSP, HSTS, CORS, etc.) for web applications to prevent common attacks. Use when user asks to "add security headers", "setup CSP", "configure CORS", "secure headers", or "HSTS setup".
backend-security-coder
sickn33
Expert in secure backend coding practices specializing in input validation, authentication, and API security. Use PROACTIVELY for backend security implementations or security code reviews.