magnus919/agent-skills/pydanticai/SKILL.md
pydanticai
Build type-safe AI agents and graph-based workflows with PydanticAI and PydanticGraph. Agent creation, function tools, capabilities, dependency injection, structured output, streaming, multi-agent patterns, testing, evals, and graph state machines. Use whenever you are building agents, tool-using LLM workflows, or graph-based state machines in Python.
- Source repository stars
- 61
- Declared platforms
- 0
- Static risk flags
- 1
- Last source update
- 2026-08-26
- Source checked
- 2026-08-28
Decision brief
What it does: where it fits
PydanticAI is a Python agent framework for building production-grade GenAI applications, built by the team behind Pydantic. PydanticGraph is its companion graph/state-machine library.
Not for
- Tasks that require unconfirmed production actions or broad system permissions.
- Environments where the pinned source and install steps cannot be inspected.
Compatibility matrix
Platform support, with evidence labels
| Platform | Status | Evidence | What to check |
|---|---|---|---|
| Codex | Not declared | No explicit evidence | Portability before use |
| Claude Code | Not declared | No explicit evidence | Portability before use |
| Cursor | Not declared | No explicit evidence | Portability before use |
| Gemini CLI | Not declared | No explicit evidence | Portability before use |
Installation
Inspect first. Install second.
The source command is displayed only when detected. A safe inspection prompt is always available so your agent can explain every action before execution.
npx skills add https://github.com/magnus919/agent-skills --skill "pydanticai"Inspect the Agent Skill "pydanticai" from https://github.com/magnus919/agent-skills/blob/531ff6753784823c878c92b988c6e55266ce09a9/pydanticai/SKILL.md at commit 531ff6753784823c878c92b988c6e55266ce09a9. List every install step, command, network request, credential, file read/write, external action, and rollback step. Explain whether it fits my task. Do not install or execute anything until I approve.
Workflow
What the source asks the agent to do
- 01
Quick Reference
python from pydanticai import Agent
python from pydanticai import Agent - 02
Basic agent — one line
agent = Agent('openai:gpt-5.2', instructions='Be concise.')
agent = Agent('openai:gpt-5.2', instructions='Be concise.') - 03
Run it
result = agent.runsync('What is the capital of France?') print(result.output) python from pydantic import BaseModel from pydanticai import Agent, RunContext
result = agent.runsync('What is the capital of France?') print(result.output) python from pydantic import BaseModel from pydanticai import Agent, RunContextclass WeatherResult(BaseModel): temperature: float conditions: stragent = Agent('openai:gpt-5.2', outputtype=WeatherResult) - 04
When to Load Which Reference
Review the “When to Load Which Reference” section in the pinned source before continuing.
Review and apply the “When to Load Which Reference” source section. - 05
Common Patterns at a Glance
→ See references/core-agents.md for full agent lifecycle, run methods, and tool patterns.
You need built-in checkpointing/persistence for long-running conversations (SQLite, Postgres backends built-in)Your multi-agent system needs subgraph composition with isolated state namespacesYou need human-in-the-loop patterns (interrupt/resume, state editing, approval workflows)
Permission review
Static risk signals and limitations
Reads files
The documentation asks the agent to read local files, directories, or repositories.
| Topic | Load When | File |Evidence record
Why each signal appears
| Signal | Value | Evidence type | Meaning |
|---|---|---|---|
| Quality score | 93/100 | Computed | Documentation, specificity, maintenance, and trust rules |
| Repository stars | 61 | Source | Repository attention, not individual Skill quality |
| Compatibility | 0 platforms | Source | Declared in the catalog source record |
| Usage guide | automated source guide | Editorial | Generated or reviewed according to the visible evidence level |
Pinned source
Provenance and original SKILL.md
- Repository
- magnus919/agent-skills
- Skill path
- pydanticai/SKILL.md
- Commit
- 531ff6753784823c878c92b988c6e55266ce09a9
- License
- MIT
- Collected
- 2026-08-28
- Default branch
- main
View the original SKILL.md
PydanticAI & PydanticGraph Expert Skill
PydanticAI is a Python agent framework for building production-grade GenAI applications, built by the team behind Pydantic. PydanticGraph is its companion graph/state-machine library.
Install:
pip install pydantic-ai # Full install (all providers)
pip install "pydantic-ai-slim[openai]" # Minimal install + your provider
Quick Reference
from pydantic_ai import Agent
# Basic agent — one line
agent = Agent('openai:gpt-5.2', instructions='Be concise.')
# Run it
result = agent.run_sync('What is the capital of France?')
print(result.output)
When to Load Which Reference
| Topic | Load When | File |
|---|---|---|
| Agent creation & lifecycle | You need to create, configure, or run an agent — define tools, deps, output types, run methods, streaming | references/core-agents.md |
| Capabilities & hooks | You need built-in capabilities (Thinking, WebSearch, MCP, etc.), on-demand loading, lifecycle hooks, or custom capabilities | references/capabilities-hooks.md |
| PydanticGraph | You need a state machine, graph-based control flow, parallel execution, BaseNode subclasses, or GraphBuilder with joins/decisions | references/graph.md |
| Models, output & streaming | You need multi-model setups, FallbackModel, streaming output, output functions, or structured output with validation | references/models-output.md |
| Multi-agent patterns & integrations | You need agent delegation, programmatic hand-off, MCP servers, durable execution, or UI adapters | references/patterns.md |
| Testing & evaluation | You need TestModel, FunctionModel, pytest patterns, overrides, or Pydantic Evals for systematic eval | references/testing-evals.md |
| Full worked examples | You want complete runnable examples — bank support agent, email feedback graph, multi-agent flight booking | references/examples.md |
| Framework boundaries | You need to compare PydanticAI vs LangGraph for a project, or want to combine them | references/hybrid-pydanticai-langgraph.md — also load skill_view(name='langgraph') |
| API surface reference | You need to find the right import path, class name, or method signature quickly | references/api-reference.md |
Common Patterns at a Glance
Agent with tools and structured output
from pydantic import BaseModel
from pydantic_ai import Agent, RunContext
class WeatherResult(BaseModel):
temperature: float
conditions: str
agent = Agent('openai:gpt-5.2', output_type=WeatherResult)
@agent.tool
async def get_weather(ctx: RunContext, city: str) -> str:
"""Get current weather for a city."""
return f"24°C and sunny in {city}"
result = agent.run_sync('Weather in London?')
print(result.output.temperature)
→ See references/core-agents.md for full agent lifecycle, run methods, and tool patterns.
Agent with dependency injection
from dataclasses import dataclass
from pydantic_ai import Agent, RunContext
@dataclass
class MyDeps:
api_key: str
db_conn: str
agent = Agent('openai:gpt-5.2', deps_type=MyDeps)
@agent.tool
async def query_db(ctx: RunContext[MyDeps], sql: str) -> str:
return f"Query results using {ctx.deps.db_conn}"
→ See references/core-agents.md for dependency injection patterns and testing overrides.
Graph with multiple nodes
from dataclasses import dataclass
from pydantic_graph import BaseNode, End, GraphRunContext, GraphBuilder
@dataclass
class MyState:
value: int = 0
@dataclass
class ProcessNode(BaseNode[MyState]):
async def run(self, ctx: GraphRunContext[MyState]) -> End | NextNode:
ctx.state.value += 1
if ctx.state.value >= 5:
return End(ctx.state.value)
return NextNode()
→ See references/graph.md for both BaseNode and GraphBuilder APIs, parallel execution, and join/reducer patterns.
When to use which run method
| When you need… | Use | Key behavior |
|---|---|---|
| A single answer, sync code | run_sync() | Blocks until complete, returns RunResult |
| A single answer, async code | run() | Async, returns RunResult |
| Stream text as it's generated | run_stream() | Async context manager, yields stream_text() / stream_output() |
| See granular events (tool calls, part starts, deltas) | run_stream_events() | Yields AgentStreamEvent types — FunctionToolCallEvent, PartStartEvent, FinalResultEvent |
| Manual control over each graph step | iter() | Iterate over agent's internal graph nodes (UserPromptNode → ModelRequestNode → CallToolsNode) |
| Tool calls to execute during streaming | run_stream_events() or run(event_stream_handler=...) | run_stream() stops at the first output that matches output_type and does NOT execute subsequent tool calls |
Details for each run method in references/core-agents.md.
Graph API: BaseNode vs GraphBuilder
| Factor | BaseNode (class-based) | GraphBuilder (function-based) |
|---|---|---|
| Style | Subclass BaseNode[StateT], implement async run() | Decorate async functions with @g.step |
| State mutation | Via ctx.state inside run() method | Via ctx.state inside step function |
| Parallelism | Manual fork/join logic | Built-in .map() per-element fan-out and .broadcast() same-input-to-multiple |
| Joins / aggregation | Manual aggregation in return types | Built-in reducers: reduce_list_append, reduce_sum, reduce_dict_update, etc. |
| Edge declaration | Inferred from run() return type annotation | Explicit via g.edge_from(source).to(target) |
| When to use | Complex node logic, OO patterns, conditional edge logic | Simple linear flows, parallel data processing, concise syntax |
Both APIs in references/graph.md.
Framework boundaries: PydanticAI vs LangGraph
Both frameworks build agentic systems with graphs and tools, but they differ sharply in design philosophy. The right choice depends on what you're optimizing for.
| Factor | PydanticAI + PydanticGraph | LangGraph | Using both together |
|---|---|---|---|
| Design philosophy | Type-safe, data-schema-driven. Feels like FastAPI. | Low-level graph primitives (Pregel/Beam inspired). Feels like NetworkX. | PydanticAI for the agent layer; LangGraph for complex orchestration |
| Agent definition | Agent(model, tools, deps, output_type) — declarative, one line | Manual StateGraph nodes with message-passing | PydanticAI Agent as a node function inside LangGraph StateGraph |
| Tool calling | @agent.tool decorator, auto-schema from type hints, RunContext DI | Manual tool registration, tool_node = ToolNode(tools) | PydanticAI's typed tool definitions used within LangGraph nodes |
| State management | GraphRunContext.state — mutable dataclass, in-memory | State with typed reducers, checkpointers (SQLite/Postgres) | LangGraph checkpointer for the outer flow; PydanticGraph for sub-graph state |
| Multi-agent patterns | Agent delegation (tool-call), programmatic hand-off, graph-based | Supervisor (central router), swarm (direct handoff), hierarchical (subgraphs) | PydanticAI delegation within a LangGraph supervisor node |
| Persistence | Durable execution via Temporal, Inngest, Prefect, DBOS | Built-in checkpointers (MemorySaver, SqliteSaver, PostgresSaver) | LangGraph checkpointer at graph level |
| Streaming | 5 methods: run, run_sync, run_stream, run_stream_events, iter | .stream() / .astream_events() on compiled graph | LangGraph .astream_events() wrapping PydanticAI event handlers |
| Learning curve | Lower — type hints guide everything | Higher — more manual wiring | Highest — two mental models |
| Best for | Single agents, tool-using workflows, type-safe structured output, teams new to agents | Complex state machines, multi-agent with branching/cycles, HITL, existing LangChain users | Large systems needing type-safe agents AND sophisticated orchestration |
Boundary conditions — consider LangGraph when:
- You need built-in checkpointing/persistence for long-running conversations (SQLite, Postgres backends built-in)
- Your multi-agent system needs subgraph composition with isolated state namespaces
- You need human-in-the-loop patterns (interrupt/resume, state editing, approval workflows)
- You're already using LangChain and want consistency
- Your graph needs cycles or dynamic fan-out via
Send()
Consider PydanticAI when:
- Type safety and IDE autocomplete are priorities
- You want declarative agents with minimal boilerplate
- You need structured output with automatic validation and retries
- Your multi-agent needs are simple delegation or sequential hand-off
- You value the composable capabilities system (Thinking, WebSearch, MCP as plugins)
Consider using both when:
- You need LangGraph's orchestration (checkpointing, subgraphs, HITL) for the outer loop, but want PydanticAI's type-safe agent definition and tool schema for the inner agent logic
- You have a mixed team: some agents benefit from PydanticAI's typing, others need LangGraph's low-level control
- See
references/hybrid-pydanticai-langgraph.mdfor a complete worked example.
LangGraph skill: skill_view(name='langgraph') — covers supervisor/swarm/hierarchical patterns, persistence, production deployment, and evals.
Error handling quick-pick
from pydantic_ai import UnexpectedModelBehavior, capture_run_messages
with capture_run_messages() as messages:
try:
result = agent.run_sync('Query')
except UnexpectedModelBehavior as e:
cause = e.__cause__ # Often ModelRetry('reason')
print(f"Root cause: {cause}")
print("Full conversation:", messages) # Inspect every message
# Common recovery: raise ModelRetry from tools with clear instructions
| Exception | Meaning | Recovery |
|---|---|---|
UnexpectedModelBehavior | Retry limit exceeded or model gave unexpected response | Inspect e.__cause__, check messages, adjust instructions or tool retries |
ModelRetry (raised from tools) | Tool wants model to retry with different args | Let it propagate — PydanticAI handles it automatically up to retries limit |
ModelAPIError | Provider returned 4xx/5xx | Check API key, rate limits, model availability |
UsageLimitExceeded | Token/request budget exhausted | Increase UsageLimits or optimize prompt |
HookTimeoutError | A lifecycle hook timed out | Increase hook timeout or optimize hook logic |
Key CLI Commands
pip install pydantic-ai # Everything
pip install "pydantic-ai-slim[openai]" # Minimal
pip install "pydantic-ai-slim[openai,google,anthropic]" # Multi-provider
Directory Structure
pydanticai/
├── SKILL.md
├── references/
│ ├── core-agents.md # Agent lifecycle, tools, deps, output
│ ├── capabilities-hooks.md # Capabilities system & lifecycle hooks
│ ├── graph.md # PydanticGraph (BaseNode + GraphBuilder)
│ ├── models-output.md # Models, streaming, structured output
│ ├── patterns.md # Multi-agent patterns & integrations
│ ├── testing-evals.md # Testing & evaluation framework
│ ├── examples.md # Complete worked examples
│ ├── hybrid-pydanticai-langgraph.md # PydanticAI as LangGraph node (hybrid pattern)
│ └── api-reference.md # Quick API surface reference
Gotchas
-
Output type = final only: The
output_typeconstrains the final response. The model can still call tools (function tools) mid-run. Output functions are different — they're forced to be called and end the run. -
pydantic-graph has zero dependency on pydantic-ai: It's a standalone library. You can use it for non-GenAI state machines. Install with
pip install pydantic-graph. -
pydantic-ai-slimvspydantic-ai: The slim package ships only core deps + OpenTelemetry. The fullpydantic-aiis a meta-package that adds openai, anthropic, google, cli, mcp, evals, web, retries, and logfire extras. -
Tool calls during streaming by default DON'T execute:
run_stream()stops at the first output that matches the output type. Userun_stream_events()orrun()withevent_stream_handlerto keep tool calls executing. -
System prompt ≠ instructions: System prompts are part of message history and round-trip. Instructions are server-side and don't appear in messages sent to clients. When reusing
message_history, the agent's new system prompt won't automatically be sent unless you addReinjectSystemPrompt. -
conversation_idis manual for forking: Passconversation_id='new'to start a fresh conversation chain from existing history. It's not automatic. -
Models named
provider:model_name— PydanticAI auto-resolves the model class from the string prefix. For custom endpoints, useOpenAIChatModel(model_name, provider=OpenAIProvider(base_url=...)). -
TestModelcan't emulate native tools: Override withagent.override(model=TestModel(), native_tools=[])in tests if your agent uses WebSearch, etc. -
defer_model_check=Truefor testable module-level agents: When declaring anAgentat module level (outside a function) and usingTestModelin tests withagent.override(model=TestModel()), setdefer_model_check=Trueon the constructor. Without it, the agent tries to resolve the model string at import time — which fails without API credentials, even though the real model is overridden before any test runs. -
Message history requires pairing: When slicing history, tool calls and their returns must stay paired or the LLM will error.
-
stream_text()fails with BaseModel output types: Whenoutput_typeis a BaseModel (structured output), callingresult.stream_text()raisesUserError('stream_text() can only be used with text responses'). Useresult.stream_output()instead to get partial validated objects as they stream in. If you need text-level streaming with structured output, userun_stream_events()and inspectPartDeltaEventwithTextPartDeltadeltas. The two methods serve different output modes — text output →stream_text(), structured output →stream_output(). -
graph.run()returns OutputT, NOT the state object: Despite passingstate=MyState()tograph.run(), the return value is the graph'soutput_type(e.g.list[int]), not the state. Thestateobject IS mutated in-place during execution (since it's a mutable dataclass), so keep a separate reference:state = MyState(items_processed=0) result = await graph.run(state=state, inputs=[1, 2, 3]) # result -> [2, 4, 6] (OutputT = list[int]) # state.items_processed -> 3 (state mutated in-place)This trap is most common with parallel
.map()patterns where the reader assumesresult.items_processedwill work. It won't. Theitems_processedcount lives on the state object you passed in, not on the return value.
Frequently asked questions
What to verify before installation and use
What does the pydanticai source document cover?
PydanticAI is a Python agent framework for building production-grade GenAI applications, built by the team behind Pydantic. PydanticGraph is its companion graph/state-machine library.
How do I install pydanticai?
The source record exposes this install command: npx skills add https://github.com/magnus919/agent-skills --skill "pydanticai". Inspect the command and pinned source before running it.
Which permission-related actions were detected?
Static rules flagged read-files in the source; the page lists the matching lines and excerpts.
Alternatives
Compare before choosing
majiayu000/spellbook
comprehensive-testing
Complete testing strategy covering TDD workflow, test pyramid, unit/integration/E2E/property testing, framework best practices (Jest, Vitest, pytest), mock strategies, and CI integration. Use when writing tests, reviewing test quality, or establishing testing standards.
rampstackco/claude-skills
data-warehouse-experimentation
Running experiments out of the data warehouse instead of via dedicated experiment platforms. SQL-based assignment, exposure logging discipline, metric definitions in dbt models, statistical analysis in SQL or Python, variance reduction with CUPED, sequential testing, and the operational tradeoffs vs platforms like Statsig and Optimizely. Triggers on warehouse-native experimentation, run experiments in BigQuery, run experiments in Snowflake, dbt experiments, SQL t-test, CUPED variance reduction,
Aperivue/medsci-skills
calc-sample-size
Interactive sample size calculator for medical research. Decision-tree guided test selection, reproducible R/Python code, effect size interpretation, and IRB-ready justification text. Supports diagnostic accuracy, agreement, proportions, continuous outcomes, survival, ANOVA, logistic regression, and non-inferiority/equivalence designs.
PramodDutta/qaskills
Pairwise Test Generator
Generate optimized test combinations using pairwise (all-pairs) testing algorithms to achieve maximum coverage with minimum test cases across multiple input parameters