LL

Automates the transformation of unstructured Telegram messages into actionable, structured knowledge atoms for Pulse Radar.

Install

mkdir -p .claude/skills/llm-pipeline && curl -L -o skill.zip "https://agentskills.codes/api/skills/download/14380" && unzip -o skill.zip -d .claude/skills/llm-pipeline && rm skill.zip

Installs to .claude/skills/llm-pipeline

Activation

This is the description your AI agent reads to decide when to run this skill — the better it matches your request, the more reliably it fires.

Pydantic-AI agents, RAG, embeddings for Pulse Radar knowledge extraction.
73 charsno explicit “when” trigger
Advanced

Key capabilities

  • Score Telegram messages for importance
  • Auto-trigger knowledge extraction based on thresholds
  • Run Pydantic AI agents for structured output
  • Save extracted knowledge to a database
  • Generate embeddings for knowledge units

How it works

The skill processes raw Telegram messages by scoring their importance, then uses Pydantic AI agents to extract structured knowledge, which is saved and embedded for pattern analysis.

Inputs & outputs

You give it
Raw Telegram messages
You get back
Structured knowledge (topics, atoms), embeddings, semantic search results

When to use llm-pipeline

  • Batch process Telegram messages into insights
  • Automate knowledge extraction from raw chat logs
  • Build AI-driven pipelines for structured data ingestion

About this skill

LLM Pipeline Skill

<overview> Pulse Radar uses LLM pipeline to transform raw Telegram messages into structured knowledge. Core philosophy: Messages individually are noise; batched extraction reveals patterns. </overview> <entity-hierarchy> ``` TOPICS (categories) └─ ATOMS (knowledge units: problem/solution/decision/insight...) └─ MESSAGES (raw data, hidden layer) ``` </entity-hierarchy> <extraction-flow> ```python # 1. Message arrives via Telegram webhook await save_telegram_message(message) # triggers TaskIQ

2. Scoring (AI Judge, not heuristics - ADR-003)

score = await importance_scorer.score(message)

classification: SIGNAL (>0.6) / NOISE (<0.3)

3. Auto-trigger extraction when threshold met

if unprocessed_count >= 10: # ai_config.message_threshold await extract_knowledge_from_messages_task.kiq()

4. KnowledgeOrchestrator runs Pydantic AI agent

agent = Agent( model=model, system_prompt=get_extraction_prompt("uk"), output_type=KnowledgeExtractionOutput, # CRITICAL: structured output output_retries=5, ) result = await agent.run(messages_content)

5. Save to DB + embed

await save_topics_and_atoms(result.output) await embed_atoms_batch_task.kiq(atom_ids)

</extraction-flow>

<agent-creation>
```python
from pydantic_ai import Agent
from pydantic_ai.models.openai import OpenAIChatModel

# Provider-specific model creation
if provider.type == "ollama":
    model = OpenAIChatModel(
        model_name=agent_config.model_name,
        provider=OllamaProvider(base_url=provider.base_url),
    )
elif provider.type == "openai":
    model = OpenAIChatModel(
        model_name=agent_config.model_name,
        provider=OpenAIProvider(api_key=api_key),
    )

# Agent with structured output
agent = Agent(
    model=model,
    output_type=MyPydanticModel,  # Forces JSON schema
    system_prompt="...",
    output_retries=5,
)
</agent-creation> <prompt-guidelines> 1. **JSON-only output** — explicitly state "respond with ONLY JSON" 2. **Schema in prompt** — include exact JSON structure expected 3. **Language enforcement** — "ALL fields MUST be in Ukrainian" 4. **Retry on language mismatch** — use `get_strengthened_prompt()` 5. **No markdown** — models often wrap JSON in ```json blocks </prompt-guidelines> <embedding-service> ```python # OpenAI: 1536 dimensions (text-embedding-3-small) # Ollama: 1024 dimensions (mxbai-embed-large) → padded to 1536

await embedding_service.generate_embedding(text) await embedding_service.embed_messages_batch(session, ids, batch_size=10)

</embedding-service>

<rag-context>
```python
# SemanticSearchService uses pgvector cosine similarity
similar_atoms = await search_service.search_atoms(
    query_embedding=embedding,
    limit=5,
    threshold=0.65,  # ai_config.semantic_search
)

# RAGContextBuilder assembles context for LLM
context = await rag_builder.build_context(
    query=user_query,
    similar_atoms=similar_atoms,
    related_messages=messages,
)
</rag-context> <context-strategies> ## RAG vs CAG
StrategyData TypePulse Radar Use
RAGDynamic (messages, atoms)Semantic search, history retrieval
CAGStatic (project config)Keywords, glossary, components preloaded

Hybrid: Project context (CAG) + similar atoms (RAG) = best extraction quality. See: @references/rag.md for detailed comparison. </context-strategies>

<adrs> - **ADR-003:** AI Importance Scoring — LLM Judge vs Heuristics (LLM chosen) - **ADR-006:** Pydantic AI vs LangChain — Hexagonal architecture (Pydantic AI chosen) </adrs> <references> - @references/architecture.md — Hexagonal LLM domain structure - @references/pydantic-ai.md — Agent configuration, streaming - @references/rag.md — RAG & CAG context strategies </references>

When not to use it

  • When individual messages are considered valuable in isolation
  • When structured output is not required
  • When a batch processing approach is not suitable

Limitations

  • Relies on a message threshold to trigger extraction
  • Requires Pydantic AI agents for structured output
  • Assumes messages individually are noise

How it compares

This skill transforms noisy individual messages into structured, embeddable knowledge through a batch-oriented, AI-driven pipeline, offering a systematic way to extract insights unlike manual message review.

Compared to similar skills

llm-pipeline side by side with the closest alternatives in the catalog.

SkillInstallsUpdatedSafetyDifficulty
llm-pipeline (this skill)02moNo flagsAdvanced
langchain268moReviewIntermediate
senior-ml-engineer67moReviewAdvanced
dspy47moReviewIntermediate

Try saying

Example prompts that trigger this skill in your AI assistant.

You might also like

langchain

zechenzhangAGI

Framework for building LLM-powered applications with agents, chains, and RAG. Supports multiple providers (OpenAI, Anthropic, Google), 500+ integrations, ReAct agents, tool calling, memory management, and vector store retrieval. Use for building chatbots, question-answering systems, autonomous agents, or RAG applications. Best for rapid prototyping and production deployments.

26138

senior-ml-engineer

davila7

World-class ML engineering skill for productionizing ML models, MLOps, and building scalable ML systems. Expertise in PyTorch, TensorFlow, model deployment, feature stores, model monitoring, and ML infrastructure. Includes LLM integration, fine-tuning, RAG systems, and agentic AI. Use when deploying ML models, building ML platforms, implementing MLOps, or integrating LLMs into production systems.

634

dspy

davila7

Build complex AI systems with declarative programming, optimize prompts automatically, create modular RAG systems and agents with DSPy - Stanford NLP's framework for systematic LM programming

430

ai-engineer

sickn33

Build production-ready LLM applications, advanced RAG systems, and intelligent agents. Implements vector search, multimodal AI, agent orchestration, and enterprise AI integrations. Use PROACTIVELY for LLM features, chatbots, AI agents, or AI-powered applications.

725

llm-app-patterns

davila7

Production-ready patterns for building LLM applications. Covers RAG pipelines, agent architectures, prompt IDEs, and LLMOps monitoring. Use when designing AI applications, implementing RAG, building agents, or setting up LLM observability.

326

llm-application-dev

skillcreatorai

Building applications with Large Language Models - prompt engineering, RAG patterns, and LLM integration. Use for AI-powered features, chatbots, or LLM-based automation.

323

Search skills

Search the agent skills registry