Provides the workflow and configuration requirements for adding support for new large language models.

Install

mkdir -p .claude/skills/dust-llm && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/4670" && unzip -o skill.zip -d .claude/skills/dust-llm && rm skill.zip

Installs to .claude/skills/dust-llm

Activation

This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.

Step-by-step guide for adding support for a new LLM in Dust. Use when adding a new model, or updating a previous one.
117 chars✓ has a “when” trigger
Intermediate

Key capabilities

  • →Configure provider-specific model IDs
  • →Map token-based pricing structures
  • →Define central model registry entries
  • →Update router whitelists
  • →Implement integration test suites

How it works

Follows a defined checklist of files across the codebase to register type definitions, pricing logic, and testing integration points for the new model.

Inputs & outputs

You give it
New model provider and technical specifications
You get back
Updated configuration files and registration registry

When to use dust-llm

  • →Register new LLM provider
  • →Update model configuration
  • →Add model tests

About this skill

Adding Support for a New LLM Model

This skill guides you through adding a newly released LLM to the model_constructors + llms stack (the endpoint-class router). It replaces the legacy lib/api/llm/clients/* router, which no longer exists.

Adding a model is usually only half the task: the model it supersedes has to be retired in the same PR, and a model the provider has switched off needs its agents repointed. See Deprecating or removing an old model.

Mental model

A model reaches production through three stacked layers. Add the new model to each:

  1. Model config (front/types/assistant/models/*) — the legacy ModelConfigurationType describing the model (context, vision, reasoning efforts, pricing tiers). Still the source of truth consumed by the UI, pricing, and the dust layer.
  2. model_constructors (front/lib/model_constructors/*) — provider-agnostic endpoint classes, one per (provider, model, region, provider-api). Each class mixes a shared provider base client with a per-model config mixin (input schema, context size, token pricing). This is where the real request/response shape and the narrowed input config live.
  3. llms (dust layer) (front/lib/llms/*) — thin Dust-specific wrappers around the model_constructors classes that add Dust concerns (display name, byok, endpoint filters, and any caps — e.g. exposing 250k context on a model that natively supports 1M). Registered into DUST_STREAM_ENDPOINTS.

Endpoints are named and filed as: {provider}_{model}_{region}_{provider_api}.ts e.g. google_gemini_3_6_flash_global_agent_platform.ts. The class name is the PascalCase of the same, with numbers spelled out: GoogleGeminiThreeDotSixFlashGlobalAgentPlatformStream.

The fastest, most reliable way to add a model is to copy the most recent model in the same family across all layers and rename. Grep every reference to that model and mirror each one. This skill lists the reference points; the sibling model is your template.

Before you start: verify against official docs (MANDATORY)

You MUST confirm every value below against the provider's official documentation and leave a URL + date in a code comment next to it. Do not carry values over from memory.

  • Specs (context window, max output tokens, vision, structured output):
    • OpenAI: https://platform.openai.com/docs/models
    • Anthropic: https://docs.anthropic.com/en/docs/about-claude/models/overview
    • Google: https://ai.google.dev/gemini-api/docs/models
    • Mistral: https://docs.mistral.ai/getting-started/models/models_overview/
  • Pricing (input / output / cached input per 1M tokens):
    • OpenAI: https://openai.com/api/pricing/
    • Anthropic: https://www.anthropic.com/pricing#anthropic-api
    • Google: https://ai.google.dev/gemini-api/docs/pricing
    • Mistral: https://mistral.ai/technology/#pricing
  • Host / region availability: verify which provider APIs and regions actually serve the model day-one. Mirror the sibling model's endpoints, but only register an endpoint whose region is actually available. Keep an unavailable-but-anticipated endpoint class defined and unregistered (see Gemini's EU agent-platform example) with a comment saying why.

WebSearch/WebFetch the docs first. If a value can't be confirmed, surface it — don't guess.

Reference points to mirror (grep the sibling model)

Pick the newest sibling (e.g. for "Gemini 3.6 Flash" the sibling is "Gemini 3.5 Flash") and grep -rln its id / const / class-name / model-id string. You will touch, roughly:

A. Model config + central registry

FileWhat to add
front/types/assistant/models/{provider}.tsX_MODEL_ID const + X_MODEL_CONFIG. Set isLatest: false on the previous model in the same family and drop "latest" from its description. Carry over the predecessor's availableIfOneOf / unavailableIfOneOf (see below).
front/types/assistant/models/models.tsAdd id to STATIC_MODEL_IDS and config to SUPPORTED_MODEL_CONFIGS (imports in both alpha blocks).
front/types/assistant/models/auto.tsIf the model should participate in auto/auto_fast/auto_complex routing, add a ModelStreamCandidate.
front/lib/model_constructors/types/models.tsAdd export const X = "model-id" and include it in the MODELS array (this is the model_constructors id type).

B. Pricing / tiers / reasoning (TYPE-ENFORCED over StaticModelIdType)

Adding the id to STATIC_MODEL_IDS makes these fail to compile until updated:

FileWhat to add
front/lib/api/assistant/token_pricing/global.tsCURRENT_MODEL_PRICING entry (input/output/cache_read_input_tokens per 1M) + doc URL comment.
front/types/assistant/models/static_model_reasoning_efforts.ts{ none, minimal, low, medium, high, xhigh, maximal } support map (satisfies Record<StaticModelIdType, ReasoningEffortSupport>). Must match the config's supportedReasoningEfforts (enforced by model_tiers.test.ts).
front/types/assistant/models/model_tiers.tsSTATIC_MODEL_TIERS entry mapping each supported effort → tier name (omit unsupported efforts).

And one that is not compile-forced, so nothing turns red if you skip it:

FileWhat to add
front/lib/api/assistant/token_pricing/eu.tsAdd the id to EU_UPLIFT_MODEL_IDS if you register a non-global endpoint that prices above its global sibling.

EU pricing is a second, silent list. Any endpoint with region = EUROPE bills through inferenceRegion: "eu" (inferenceRegionForEndpointRegion in front/lib/api/llm/transitionLLM.ts), and computeTokensCostForUsageInMicroUsd then looks the model up in EU_MODEL_PRICING — falling back to the global rate when it is absent. EU_UPLIFT_MODEL_IDS is satisfies readonly StaticModelIdType[], which validates the ids present but does not force completeness, so a missing entry undercharges EU traffic forever with nothing failing.

The uplift is per provider and per endpoint, not per model — compare the two endpoint classes' tokenPricing rather than assuming. Regional agent-platform (Vertex) endpoints charge 10% over global for both Anthropic and Google, so a new Gemini registered on eu/agent-platform belongs in the list just as much as a Claude does. OpenAI uplifts only the models whose pricing page lists a data-residency premium (gpt-5.4/5.5/5.6/6 yes, gpt-5/5.1/5.2 no). Mistral's own models are EU-only with no global sibling, so nothing to add for them.

EU_MODEL_PRICING derives every field by multiplying the global entry by EU_PRICING_MULTIPLIER, so it is only correct when the EU endpoint is a flat 1.1× of global. A non-uniform regional price needs an explicit entry in EU_HOST_MODEL_PRICING, not the multiplier — always the case when the EU host differs from the global one (GLM-5.3: Fireworks global, Mistral EU).

Third-party model on a lab's own host (e.g. GLM-5.3 on Mistral): set lab on the endpoint (a host serving several labs leaves it off its base client), map the host's model name with modelToHostModel, and check PROVIDER_ID_TO_HOST in front/lib/api/llm/index.ts. Routing matches lab ∈ whitelisted labs OR host ∈ whitelisted hosts, so an endpoint whose lab and host are both unmapped is silently unreachable.

Gating is inherited, and lives in two unlinked places. A new version of a gated model stays gated — being newer is not a reason to release it. Copy the predecessor's availableIfOneOf / unavailableIfOneOf onto the new X_MODEL_CONFIG (gates the picker, via isModelAvailable) and declare the same flag on every endpoint you add (gates the router, via isEndpointAvailable):

static readonly endpointFilter = {
  featureFlags: { contains: "fireworks_new_model_feature" as const },
};

Half-gating fails silently either way: hidden but reachable, or pickable but unroutable — and resolveModel swaps in a fallback model instead of erroring. Releasing a gated family is a separate, deliberate change.

An eu/agent-platform endpoint takes EU_AGENT_PLATFORM_ENDPOINT_FILTER, never {}. Dust-managed EU hosting is sold to credit-priced plans and to workspaces carrying use_vertex_for_supported_models; that rule lives once, in front/lib/llms/utils/endpoint_filters.ts, and every *_eu_agent_platform.ts dust wrapper assigns it verbatim:

static readonly endpointFilter = EU_AGENT_PLATFORM_ENDPOINT_FILTER;

Copy the constant, not a sibling's inline literal — a hand-written copy is how six Gemini Flash EU endpoints ended up on {}, routing legacy workspaces to EU hosting nobody had promised them while the picker showed no EU flag. The client mirrors the same constant in useRunsOnRegionalHosting (front/hooks/useRunsOnRegionalHosting.ts) to decide whether to show a workspace its hosting region, so an endpoint that opts out makes that indicator lie.

A model-specific gate composes with it rather than replacing it — { and: [FILTER, { featureFlags: … }] }. This is about region = EUROPE on host = AGENT_PLATFORM only: provider-hosted EU endpoints (*_eu_openai_responses, *_eu_mistral) run on the provider's own EU infrastructure, are available to everyone, and keep endpointFilter = {}.

C. model_constructors — the endpoint classes (stream)

FileWhat to add
front/lib/model_constructors/providers/{provider}/models/{model}.tsConfig mixin WithXConfig(Base) exposing static model, static configSchema, static contextSize, static maxOutputTokens. Reuse the provider's shared inputConfig/reasoning_efforts/shared helpers. **contextSize/maxOutputTokens are the REAL provider valu

Content truncated.

When not to use it

  • →When only updating system prompts without changing model infra
  • →When the model is already registered in the registry
  • →If adding support for a provider not supported by Dust architecture

Prerequisites

Access to Dust front-end repositoryModel specifications (context size, pricing, tokenizers)

Limitations

  • →Requires manual verification of provider docs
  • →Config files are highly specific to the Dust architecture
  • →Manual test update required

How it compares

It provides a structured schema for integration rather than relying on trial-and-error configuration updates.

Compared to similar skills

dust-llm side by side with the closest alternatives in the catalog.

SkillInstallsUpdatedSafetyDifficulty
dust-llm (this skill)13moReviewIntermediate
langchain2610moReviewIntermediate
ai-sdk114moReviewAdvanced
langfuse78moNo flagsIntermediate

Try saying

Example prompts that trigger this skill in your AI assistant.

You might also like

langchain

zechenzhangAGI

Framework for building LLM-powered applications with agents, chains, and RAG. Supports multiple providers (OpenAI, Anthropic, Google), 500+ integrations, ReAct agents, tool calling, memory management, and vector store retrieval. Use for building chatbots, question-answering systems, autonomous agents, or RAG applications. Best for rapid prototyping and production deployments.

26138

ai-sdk

vercel

Answer questions about the AI SDK and help build AI-powered features. Use when developers: (1) Ask about AI SDK functions like generateText, streamText, ToolLoopAgent, embed, or tools, (2) Want to build AI agents, chatbots, RAG systems, or text generation features, (3) Have questions about AI providers (OpenAI, Anthropic, Google, etc.), streaming, tool calling, structured output, or embeddings, (4) Use React hooks like useChat or useCompletion. Triggers on: "AI SDK", "Vercel AI SDK", "generateText", "streamText", "add AI to my app", "build an agent", "tool calling", "structured output", "useChat".

1150

langfuse

davila7

Expert in Langfuse - the open-source LLM observability platform. Covers tracing, prompt management, evaluation, datasets, and integration with LangChain, LlamaIndex, and OpenAI. Essential for debugging, monitoring, and improving LLM applications in production. Use when: langfuse, llm observability, llm tracing, prompt management, llm evaluation.

743

llm-application-dev

skillcreatorai

Building applications with Large Language Models - prompt engineering, RAG patterns, and LLM integration. Use for AI-powered features, chatbots, or LLM-based automation.

323

llm-patterns

alinaqi

AI-first application patterns, LLM testing, prompt management

10

azure-ai-projects-ts

microsoft

Build AI applications using Azure AI Projects SDK for JavaScript (@azure/ai-projects). Use when working with Foundry project clients, agents, connections, deployments, datasets, indexes, evaluations, or getting OpenAI clients.

00

Search skills

Search the agent skills registry