EX

exa-performance-tuning

Optimizes Exa API latency by selecting search types based on requirements and implementing caching.

Install

mkdir -p .claude/skills/exa-performance-tuning && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/9311" && unzip -o skill.zip -d .claude/skills/exa-performance-tuning && rm skill.zip

Installs to .claude/skills/exa-performance-tuning

Activation

This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.

Optimize Exa API performance with search type selection, caching, and
69 charsno explicit “when” trigger
Intermediate

Key capabilities

  • Select an Exa search type based on latency requirements
  • Minimize content retrieval to reduce latency
  • Cache search results using an LRU cache
  • Parallelize independent Exa search queries
  • Implement a two-phase search for selective content retrieval
  • Normalize queries to increase cache hit rates

How it works

The skill optimizes Exa API performance by selecting appropriate search types, reducing result content, caching responses, and executing queries in parallel.

Inputs & outputs

You give it
Exa search query and desired latency budget in milliseconds
You get back
Optimized Exa search results with reduced latency and improved throughput

When to use exa-performance-tuning

  • Improve search response times
  • Implement request caching
  • Optimize production search workloads

About this skill

Exa Performance Tuning

Overview

Optimize Exa search API response times for production workloads. Key levers: search type selection (instant < fast < auto < neural < deep), result count reduction, content scope control, result caching, and parallel query execution.

Latency by Search Type

TypeTypical LatencyUse Case
instant< 150msReal-time autocomplete, typeahead
fastp50 < 425msSpeed-critical user-facing search
auto300-1500msGeneral purpose (default)
neural500-2000msBest semantic quality
deep2-5sMaximum coverage, light deep search
deep-reasoning5-15sComplex research questions

Instructions

Step 1: Match Search Type to Latency Budget

import Exa from "exa-js";

const exa = new Exa(process.env.EXA_API_KEY);

function selectSearchType(latencyBudgetMs: number) {
  if (latencyBudgetMs < 200) return "instant";
  if (latencyBudgetMs < 500) return "fast";
  if (latencyBudgetMs < 1500) return "auto";
  if (latencyBudgetMs < 3000) return "neural";
  return "deep";
}

async function optimizedSearch(query: string, latencyBudgetMs: number) {
  const type = selectSearchType(latencyBudgetMs);
  const numResults = latencyBudgetMs < 500 ? 3 : latencyBudgetMs < 2000 ? 5 : 10;

  return exa.search(query, { type, numResults });
}

Step 2: Minimize Content Retrieval

// Each content option adds latency. Only request what you need.

// Fastest: metadata only (no content retrieval)
const metadataOnly = await exa.search("query", { numResults: 5 });

// Medium: highlights only (much smaller than full text)
const highlightsOnly = await exa.searchAndContents("query", {
  numResults: 5,
  highlights: { maxCharacters: 300 },
  // No text or summary — saves content retrieval time
});

// Slower: full text (use maxCharacters to limit)
const withText = await exa.searchAndContents("query", {
  numResults: 3,  // fewer results = faster
  text: { maxCharacters: 1000 },  // limit content size
});

Step 3: Cache Search Results

import { LRUCache } from "lru-cache";

const searchCache = new LRUCache<string, any>({
  max: 5000,
  ttl: 2 * 3600 * 1000, // 2-hour TTL
});

async function cachedSearch(query: string, opts: any) {
  const key = `${query}:${opts.type || "auto"}:${opts.numResults || 10}`;
  const cached = searchCache.get(key);
  if (cached) return cached; // Cache hit: 0ms vs 500-2000ms

  const results = await exa.search(query, opts);
  searchCache.set(key, results);
  return results;
}

Step 4: Parallelize Independent Searches

// Run independent queries concurrently instead of sequentially
async function parallelSearch(queries: string[]) {
  const searches = queries.map(q =>
    cachedSearch(q, { type: "auto", numResults: 3 })
  );
  return Promise.all(searches);
  // 3 parallel searches: ~600ms total (limited by slowest)
  // 3 sequential searches: ~1800ms total
}

Step 5: Two-Phase Search Pattern

// Phase 1: Fast search for URLs only
// Phase 2: Selective content retrieval for top results only
async function twoPhaseSearch(query: string) {
  // Phase 1: metadata only (fast)
  const results = await exa.search(query, { type: "auto", numResults: 10 });

  // Phase 2: get content only for top 3 results
  const topUrls = results.results.slice(0, 3).map(r => r.url);
  const contents = await exa.getContents(topUrls, {
    text: { maxCharacters: 2000 },
    highlights: { maxCharacters: 500, query },
  });

  return contents;
  // Saves content retrieval time for 7 results you won't use
}

Step 6: Query Normalization for Cache Hits

function normalizeQuery(query: string): string {
  return query
    .toLowerCase()
    .trim()
    .replace(/\s+/g, " ")       // collapse whitespace
    .replace(/[?.!,;:]+$/, ""); // strip trailing punctuation
}

async function normalizedSearch(query: string, opts: any) {
  return cachedSearch(normalizeQuery(query), opts);
}
// Increases cache hit rate by 20-40% for user-generated queries

Performance Comparison

StrategyLatency SavingsImplementation
instant type5-10x faster than neuralOne-line change
Reduce numResults (10 -> 3)~200-500ms savedOne-line change
Highlights instead of text~100-300ms savedReplace text with highlights
LRU cache100% for cache hits~20 lines
Parallel queries2-3x throughputPromise.all wrapper
Two-phase search~30-50% for large result sets~15 lines

Error Handling

IssueCauseSolution
Search taking 3s+Neural search on complex querySwitch to fast or auto type
Timeout on contentLarge pages, slow sourcesSet maxCharacters limit
Cache miss rate highUnique queries each timeNormalize queries before caching
Rate limit (429)Too many concurrent searchesAdd request queue with concurrency limit

Resources

Next Steps

For cost optimization, see exa-cost-tuning. For reliability, see exa-reliability-patterns.

When not to use it

  • When maximum coverage is required, as this may increase latency
  • When complex research questions require deep-reasoning search type

Limitations

  • Neural search on complex queries can take 3 seconds or more
  • Large pages or slow sources can cause content retrieval timeouts
  • Unique queries each time can lead to a high cache miss rate

How it compares

This skill provides specific strategies like two-phase search and query normalization to improve Exa API performance, rather than just making basic API calls.

Compared to similar skills

exa-performance-tuning side by side with the closest alternatives in the catalog.

SkillInstallsUpdatedSafetyDifficulty
exa-performance-tuning (this skill)327dReviewIntermediate
mcp-builder1363moReviewAdvanced
stripe-integration482moNo flagsAdvanced
copilot-sdk74moReviewIntermediate

Try saying

Example prompts that trigger this skill in your AI assistant.

More by jeremylongshore

View all by jeremylongshore

analyzing-logs

jeremylongshore

Analyze application logs to detect performance issues, identify error patterns, and improve stability by extracting key insights.

14123

ollama-setup

jeremylongshore

Configure auto-configure Ollama when user needs local LLM deployment, free AI alternatives, or wants to eliminate hosted API costs. Trigger phrases: "install ollama", "local AI", "free LLM", "self-hosted AI", "replace OpenAI", "no API costs". Use when appropriate context detected. Trigger with relevant phrases based on skill purpose.

1167

backtesting-trading-strategies

jeremylongshore

Backtest crypto and traditional trading strategies against historical data. Calculates performance metrics (Sharpe, Sortino, max drawdown), generates equity curves, and optimizes strategy parameters. Use when user wants to test a trading strategy, validate signals, or compare approaches. Trigger with phrases like "backtest strategy", "test trading strategy", "historical performance", "simulate trades", "optimize parameters", or "validate signals".

1071

generating-database-seed-data

jeremylongshore

Process this skill enables AI assistant to generate realistic test data and database seed scripts for development and testing environments. it uses faker libraries to create realistic data, maintains relational integrity, and allows configurable data volumes. u... Use when working with databases or data models. Trigger with phrases like 'database', 'query', or 'schema'.

1033

cursor-codebase-indexing

jeremylongshore

Execute set up and optimize Cursor codebase indexing. Triggers on "cursor index setup", "codebase indexing", "index codebase", "cursor semantic search". Use when working with cursor codebase indexing functionality. Trigger with phrases like "cursor codebase indexing", "cursor indexing", "cursor".

885

testing-mobile-apps

jeremylongshore

Execute mobile app testing on iOS and Android devices/simulators. Use when performing specialized testing. Trigger with phrases like "test mobile app", "run iOS tests", or "validate Android functionality".

810

You might also like

mcp-builder

anthropics

Guide for creating high-quality MCP (Model Context Protocol) servers that enable LLMs to interact with external services through well-designed tools. Use when building MCP servers to integrate external APIs or services, whether in Python (FastMCP) or Node/TypeScript (MCP SDK).

136215

stripe-integration

wshobson

Implement Stripe payment processing for robust, PCI-compliant payment flows including checkout, subscriptions, and webhooks. Use when integrating Stripe payments, building subscription systems, or implementing secure checkout flows.

48165

copilot-sdk

github

Build agentic applications with GitHub Copilot SDK. Use when embedding AI agents in apps, creating custom tools, implementing streaming responses, managing sessions, connecting to MCP servers, or creating custom agents. Triggers on Copilot SDK, GitHub SDK, agentic app, embed Copilot, programmable agent, MCP server, custom agent.

763

openrouter-hello-world

jeremylongshore

Create your first OpenRouter API request with a simple example. Use when learning OpenRouter or testing your setup. Trigger with phrases like 'openrouter hello world', 'openrouter first request', 'openrouter quickstart', 'test openrouter'.

733

telegram-dev

2025Emma

Telegram 生态开发全栈指南 - 涵盖 Bot API、Mini Apps (Web Apps)、MTProto 客户端开发。包括消息处理、支付、内联模式、Webhook、认证、存储、传感器 API 等完整开发资源。

232

mistral-upgrade-migration

jeremylongshore

Analyze, plan, and execute Mistral AI SDK upgrades with breaking change detection. Use when upgrading Mistral SDK versions, detecting deprecations, or migrating to new API versions. Trigger with phrases like "upgrade mistral", "mistral migration", "mistral breaking changes", "update mistral SDK", "analyze mistral version".

17

Search skills

Search the agent skills registry