yara-rule-authoring
Guidance and optimization for writing YARA-X malware detection rules.
Install
mkdir -p .claude/skills/yara-rule-authoring && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/4616" && unzip -o skill.zip -d .claude/skills/yara-rule-authoring && rm skill.zipInstalls to .claude/skills/yara-rule-authoring
Activation
This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.
Guides authoring of high-quality YARA-X detection rules for malware identification. Use when writing, reviewing, or optimizing YARA rules. Covers naming conventions, string selection, performance optimization, migration from legacy YARA, and false positive reduction. Triggers on: YARA, YARA-X, malware detection, threat hunting, IOC, signature, crx module, dex module.Key capabilities
- →Author YARA-X detection rules
- →Optimize rule performance
- →Migrate legacy YARA rules
- →Reduce false positives
How it works
Guides the creation of YARA-X rules using atom-based string selection and performance-optimized conditions.
Inputs & outputs
When to use yara-rule-authoring
- →Write a new YARA-X rule
- →Optimize YARA rule performance
- →Migrate legacy YARA to YARA-X
About this skill
YARA-X Rule Authoring
Write detection rules that catch malware without drowning in false positives.
This skill targets YARA-X, the Rust-based successor to legacy YARA — 5-10x faster regex, better errors, built-in formatter, stricter validation, new modules (crx, dex), 99% rule compatibility. It powers VirusTotal's production systems. Install with brew install yara-x or cargo install yara-x; the CLI is yr. See Migrating from Legacy YARA for existing rules.
Core Principles
-
Strings must generate good atoms — YARA extracts 4-byte subsequences for fast matching. Strings with repeated bytes, common sequences, or under 4 bytes force slow bytecode verification on too many files.
-
Target specific families, not categories — "Detects ransomware" catches everything and nothing. "Detects LockBit 3.0 configuration extraction routine" catches what you want.
-
Test against goodware before deployment — A rule that fires on Windows system files is useless. Validate against VirusTotal's goodware corpus or your own clean file set.
-
Short-circuit with cheap checks first —
filesize(instant), then magic bytes (nearly instant), then strings (cheap), then modules (expensive). -
Metadata is documentation — Future you (and your team) need to know what this catches, why, and where the sample came from.
When to Use
- Writing new YARA-X rules for malware detection
- Reviewing existing rules for quality or performance issues
- Optimizing slow-running rulesets
- Converting IOCs or threat intel into detection signatures
- Debugging false positive issues
- Preparing rules for production deployment
- Migrating legacy YARA rules to YARA-X
- Analyzing Chrome extensions (crx module) or Android apps (dex module)
When NOT to Use
- Static analysis requiring disassembly → use Ghidra/IDA skills
- Dynamic malware analysis → use sandbox analysis skills
- Network-based detection → use Suricata/Snort skills
- Memory forensics with Volatility → use memory forensics skills
- Simple hash-based detection → just use hash lists
Platform Considerations
YARA works on any file type. Adapt patterns to your target:
| Platform | Magic Bytes | Bad Strings | Good Strings |
|---|---|---|---|
| Windows PE | uint16(0) == 0x5A4D | API names, Windows paths | Mutex names, PDB paths |
| macOS Mach-O | uint32(0) == 0xFEEDFACE (32-bit), 0xFEEDFACF (64-bit), uint32be(0) == 0xCAFEBABE (universal) | Common Obj-C methods | Keylogger strings, persistence paths |
| JavaScript/Node | (none needed) | require, fetch, axios | Obfuscator signatures, eval+decode chains |
| npm/pip packages | (none needed) | postinstall, dependencies | Suspicious package names, exfil URLs |
| Office docs | uint32(0) == 0x04034B50 | VBA keywords | Macro auto-exec, encoded payloads |
| VS Code extensions | (none needed) | vscode.workspace | Uncommon activationEvents, hidden file access |
| Chrome extensions | Use crx module | Common Chrome APIs | Permission abuse, manifest anomalies |
| Android apps | Use dex module | Standard DEX structure | Obfuscated classes, suspicious permissions |
uintNN()reads little-endian. Write the constant as the bytes reversed, or useuintNNbe()and write them in file order. A ZIP/OOXML file starts with bytes50 4B 03 04, so it isuint32(0) == 0x04034B50—uint32(0) == 0x504B0304compiles cleanly and never matches anything. The same trap catches Mach-O universal binaries: on disk they areCA FE BA BE, souint32(0) == 0xCAFEBABEis a dead branch; writeuint32be(0) == 0xCAFEBABEoruint32(0) == 0xBEBAFECA. Verify withyr scanagainst one known-good sample before trusting any magic-byte check.
macOS Malware Detection
No dedicated Mach-O module exists yet — use magic bytes plus string patterns. Good indicators:
- Keylogger artifacts:
CGEventTapCreate,kCGEventKeyDown - SSH tunnel strings:
ssh -D,tunnel,socks - Persistence paths:
~/Library/LaunchAgents,/Library/LaunchDaemons - Credential theft:
security find-generic-password,keychain
// Pattern from Airbnb BinaryAlert
rule SUSP_Mac_ProtonRAT
{
strings:
$lib1 = "SRWebSocket" ascii // Library indicators
$lib2 = "SocketRocket" ascii
$behav1 = "SSH tunnel not launched" ascii // Behavioral indicators
$behav2 = "Keylogger" ascii
condition:
(uint32(0) == 0xFEEDFACF or uint32be(0) == 0xCAFEBABE) and
any of ($lib*) and any of ($behav*)
}
JavaScript Detection
| Target | Approach |
|---|---|
| npm package | package.json patterns, postinstall/preinstall hooks, exfil combination: fetch + env access + credential paths |
| Chrome extension | crx module |
| Other extension | Manifest patterns, background script behaviors |
| Standalone JS | Obfuscation markers (eval+atob, fromCharCode chains), unique function/variable names, packed payloads |
| Minified/webpack bundle | Unique strings that survive bundling (URLs, magic values); avoid function names — they get mangled |
Good JS strings: Ethereum function selectors — { a9 05 9c bb } (transfer(address,uint256)), { 70 a0 82 31 } (balanceOf(address)); zero-width characters for steganography — { E2 80 8B E2 80 8C }; obfuscator signatures — _0x, var _0x; specific C2 domains and webhook URLs.
Bad JS strings: require, fetch, axios (too common); Buffer, crypto (legitimate uses everywhere); process.env alone (need specific env var names).
String Selection
Value ranking: mutex names are gold, C2 paths silver, error messages bronze. Stack strings are almost always unique. If you need more than 6 strings, you're over-fitting.
Reject a candidate string when any of these holds:
| Test | Why it fails | Do instead |
|---|---|---|
| Under 4 bytes | No atom | Find a longer string |
Repeated bytes (0000, 9090) | Weak atom | Add surrounding context |
API name (VirtualAlloc, CreateRemoteThread) | Every packer and installer calls it | Hex pattern of the call site plus a unique marker |
| Appears in Windows system files | Guaranteed FPs | Find something family-specific |
Common path (C:\Windows\, cmd.exe) | Ubiquitous | Find malware-specific paths |
| Appears in other malware families | Not identifying this family | Combine with a family-specific marker |
Everything left — unique to this family — is what the rule should rest on.
Choosing a String Type
| Need | Use |
|---|---|
| Exact ASCII/Unicode text | $s = "MutexName" ascii wide |
| Specific byte sequence | $h = { 4D 5A 90 00 } |
| Byte sequence with variation | Hex wildcards: { 4D 5A ?? ?? 50 45 } |
| Pattern with structure (URLs, paths) | Bounded regex: /https:\/\/[a-z]{5,20}\.onion/ |
| Unknown encoding (XOR, base64) | Modifier: $s = "config" xor(0x00-0xFF) |
Modifier discipline: never use nocase or wide speculatively — only with confirmed evidence that case or encoding varies across samples. nocase doubles atom generation; wide doubles string matching. "If you don't have a clear reason for using those modifiers, don't do it" — Kaspersky Applied YARA.
Condition Design
Order for short-circuit: filesize <, magic bytes, strings, modules. If the condition runs past 5 lines, split into multiple rules.
all of vs any of
| Situation | Use |
|---|---|
| Strings are individually unique to the malware | any of them — each alone is suspicious |
| Strings are common but the combination is suspicious | all of them — require the full pattern |
| Strings have different confidence levels | Group: all of ($core_*) and any of ($variant_*) |
| Seeing false positives | Tighten: any → all, add more required strings |
Lesson from production: rules using any of ($network_*) where the strings included fetch, axios, and http matched virtually all web applications. Switching to require a credential path AND a network call AND an exfil destination eliminated the FPs.
Grouping by Confidence
Different indicator types carry different weight — a C2 domain might be definitive while library imports need corroboration. Grouping by prefix lets you express graduated requirements:
strings:
$a1 = "SRWebSocket" ascii // Category A: library indicators
$a2 = "SocketRocket" ascii
$b1 = "SSH tunnel" ascii // Category B: behavioral
$b2 = "keylogger" ascii nocase
$c1 = /https:\/\/[a-z0-9]{8,16}\.onion/ // Category C: C2
condition:
filesize < 10MB and
any of ($a*) and any of ($b*) // Evidence from BOTH categories
Modules vs Byte Checks
| Need | Use |
|---|---|
| imphash, rich header, authenticode | PE module — too complex to replicate |
| Magic bytes or simple offsets | uint16/uint32 — faster, no module overhead |
| Section names/sizes | PE module, but put the magic-byte filter FIRST |
| Chrome extension permissions | crx module — string parsing is fragile |
| LNK target paths | lnk module — the format is complex |
"Avoid the magic module — use explicit hex checks instead" — Neo23x0. Generalize it: if uint32() can do the job, don't load a module.
Performance
- Regex must be anchored to a 4+ byte literal. Without one it evaluates at every file offset — catastrophic. Write
/mshta\.exe http:\/\/.../, not/http:\/\/.../. If you can't anchor, use a hex pattern with wildcards. - Bound every regex quantifier —
.{0,30}, never.*. Unbounded regex is both a performance disaster and a memory explosion. - Bound loops with filesize —
filesize < 100KB and for all i in (1..#a) : .... Unbounded#acan reach thousands in large files. - Prefer hex over regex where the bytes are fixed.
Before Writing: Is the Sample Packed?
| Signal | What to do |
|---|---|
| Entropy > 7.0 | Likely packed — find the unpacked layer first |
| F |
Content truncated.
When not to use it
- →Static analysis
- →Dynamic malware analysis
- →Network-based detection
Prerequisites
Limitations
- →Requires YARA-X compatible syntax
How it compares
Focuses on YARA-X specific optimizations and performance guidelines rather than generic YARA rules.
Compared to similar skills
yara-rule-authoring side by side with the closest alternatives in the catalog.
| Skill | Installs | Updated | Safety | Difficulty |
|---|---|---|---|---|
| yara-rule-authoring (this skill) | 1 | 2mo | Caution | Advanced |
| protocol-reverse-engineering | 9 | 6mo | Review | Advanced |
| equilateral-agents | 5 | 9mo | No flags | Intermediate |
| secops-triage | 4 | 7mo | No flags | Intermediate |
Try saying
Example prompts that trigger this skill in your AI assistant.
More by trailofbits
View all by trailofbits →You might also like
protocol-reverse-engineering
wshobson
Master network protocol reverse engineering including packet analysis, protocol dissection, and custom protocol documentation. Use when analyzing network traffic, understanding proprietary protocols, or debugging network communication.
equilateral-agents
Equilateral-AI
22 production-ready AI agents with database-driven orchestration for security reviews, code quality analysis, deployment validation, infrastructure checks, and compliance. Auto-activates for security concerns, deployment tasks, code reviews, quality checks, and compliance questions. Includes upgrade paths to enterprise features (GDPR, HIPAA, multi-account AWS, ML-based optimization).
secops-triage
Expert guidance for security alert triage. Use this when the user asks to "triage" an alert or case.
netflows
BrownFineSecurity
Network flow extractor that analyzes pcap/pcapng files to identify outbound connections with automatic DNS hostname resolution. Use when you need to enumerate network destinations, identify what hosts a device communicates with, or map IP addresses to hostnames from packet captures.
azure-bgp
benchflow-ai
Analyze and resolve BGP oscillation and BGP route leaks in Azure Virtual WAN–style hub-and-spoke topologies (and similar cloud-managed BGP environments). Detect preference cycles, identify valley-free violations, and propose allowed policy-level mitigations while rejecting prohibited fixes.
secops-investigate
Expert guidance for deep security investigations. Use this when the user asks to "investigate" a case, entity, or incident.