Creator · existential-birds
Last updated · Sep 7, 2026
Test PydanticAI agents using TestModel, FunctionModel, VCR cassettes, and inline snapshots. Use when writing unit tests, mocking LLM responses, or recording API interactions.
Sandbox only
Install targets
Codex install prompt
Install the "pydantic-ai-testing" agent skill from https://github.com/existential-birds/beagle/tree/main/plugins/beagle-ai/skills/pydantic-ai-testing. Read its SKILL.md or equivalent instructions first, install only the files needed for this workspace, and summarize any required setup before using it. Skill purpose: Test PydanticAI agents using TestModel, FunctionModel, VCR cassettes, and inline snapshots. Use when writing unit tests, mocking LLM responses, or recording API interactions. After the install attempt, report the result to https://www.openagentskill.com/api/agent/outcome with POST JSON {"event_id":"install_<unique-id>","skill_slug":"existential-birds-pydantic-ai-testing","task":"Install pydantic-ai-testing","agent":"codex","outcome":"success","install_used":true}. Replace event_id with a unique value and outcome with success or failed. Report success only after the skill is installed and a minimal verification passes.Supply asset profile
Code review, repo analysis, testing, CI, GitHub, DevOps, and developer workflow skills.
Scenario
GitHub automation
I need my agent to triage GitHub issues, review pull requests, and summarize repository changes.
Agent fit
Claude Code + OpenAI Agents + CLI
Codex, Claude Code, Cursor, CLI, or custom agents.
Install
Ready
npx skills add existential-birds/beagle --skill pydantic-ai-testing
Maintenance
fresh
29d since push
Risk
Needs review
Quality score needs review
GitHub quality
80
66/100 Quality · 73/100 Trust
Coverage tags
Review notes
Quality score needs review · GitHub adoption: 80 GitHub stars
Agent adoption scorecard
These scores combine public repository metadata, OpenAgentSkill review signals, maintenance freshness, and install readiness. They are a shortlist signal, not a replacement for human review.
Quality
PromisingUseful candidate, but compare it with alternatives before adopting.
Trust
Sandbox onlyUseful candidate with missing or mixed trust signals. Keep it in an isolated workspace until the outcome loop proves task fit.
Audit
Needs reviewA machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
OpenAgentSkill Trust Score v5
Run only in a sandbox and compare close alternatives before using it for real work.
Stars
80 GitHub stars
Repo activity
80 stars, 8 forks
Maintenance
29d since push
License
Apache-2.0
Install
npx skills add existential-birds/beagle --skill pydantic-ai-testing
Install safety
Agent-readable metadata
Use this block or the embedded JSON to decide whether an agent should install this skill, choose an alternative, or ask for human review first.
Suited tasks
Suited agents
Install decision
Trust and risk
Outcome loop
Install command
npx skills add existential-birds/beagle --skill pydantic-ai-testingDo not use when
Alternative
168.6K Stars
npx skills add mattpocock/skills --skill code-review
Alternative
40.8K Stars
npx skills add appsmithorg/appsmith
Alternative
175.7K Stars
npx skills add mattpocock/skills --skill implement
Alternative
30.9K Stars
npx skills add vercel-labs/agent-skills --skill vercel-react-best-practices
Agent safety v2
Sparse or mixed signals. Useful for discovery, but not for autonomous installation.
Test manually in an isolated workspace and compare against safer alternatives.
high
Skill metadata references terminal, CLI, shell, subprocess, or command execution workflows.
medium
Skill likely fetches remote pages, APIs, repositories, or external services.
medium
Skill may read or write project files, documents, generated artifacts, or local workspace state.
Agent resolve plan
The Resolve API returns the selected skill, alternatives, safety policy, audit notes, install target, and copy-paste prompt an agent can follow without scraping this page.
Open JSON
/api/agent/resolve?task=Use%20pydantic-ai-testing%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Resolve text
/api/agent/resolve?task=Use%20pydantic-ai-testing%20for%20an%20agent%20workflow&agent=codex&max_risk=medium&format=text
Install handoff
/api/skills/existential-birds-pydantic-ai-testing/install
Agent should check
Copy prompt
Task: Use pydantic-ai-testing in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20pydantic-ai-testing%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/existential-birds-pydantic-ai-testing/install
Install command: npx skills add existential-birds/beagle --skill pydantic-ai-testing
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.Agent handoff
Use the public install endpoint to fetch the command, safety checklist, target prompts, and canonical links for this skill.
Install handoff
/api/skills/existential-birds-pydantic-ai-testing/install
LLM text format
/api/skills/existential-birds-pydantic-ai-testing/install?format=text
Find alternatives
/api/skills/search?q=pydantic-ai-testing&limit=3
Agent prompt
Use pydantic-ai-testing for this task. Review https://www.openagentskill.com/api/skills/existential-birds-pydantic-ai-testing/install, then install with: npx skills add existential-birds/beagle --skill pydantic-ai-testingRegistry metadata
This page exposes the same decision, trust, audit, use-case, and install signals through the Registry API, so agents can rank this skill without scraping the UI.
Manifest
/api/registry/manifest/existential-birds-pydantic-ai-testing
LLM text
/api/registry/manifest/existential-birds-pydantic-ai-testing?format=text
Install alias
/api/registry/install/existential-birds-pydantic-ai-testing
Recommend
/api/registry/recommend?task=Use%20pydantic-ai-testing%20in%20an%20agent%20workflow&limit=3
Agent fit
GitHub automation
Use-case tags
Platforms
Claude Code, OpenAI Agents
Audit report
A machine-readable review of install readiness, security metadata, maintenance, and adoption risk.
Agent decision cockpit
Prototype with this skill first; keep a fallback candidate ready.
Role in stack
Fallback candidate
Primary fit
GitHub automation
Trust label
Prototype first
Install path
Command ready
Use when
Evidence
review first
Implementation path
Trust profile
Useful candidate with missing or mixed trust signals. Keep it in an isolated workspace until the outcome loop proves task fit.
GitHub adoption
CHECK80 GitHub stars
Stars/forks activity
CHECK80 stars, 8 forks; issue activity unavailable in current metadata
Recent maintenance
PASS29d since push
License clarity
PASSApache-2.0
Good signals
Review before install
Recommended action
Run only in a sandbox and compare close alternatives before using it for real work.
Quality profile
Useful candidate, but compare it with alternatives before adopting.
Workflow fit
Manage repositories
I need my agent to triage GitHub issues, review pull requests, and summarize repository changes.
Verify behavior
I need my agent to test a web app, reproduce bugs, and verify fixes.
Operate local tools
I need my agent to operate local files and desktop apps in a repeatable workflow.
Workflow fit
Inspect, patch, and verify code
A workflow for software agents that inspect repositories, review pull requests, generate tests, and turn findings into shippable patches.
Turn skills into distribution
A workflow for turning newly indexed skills into SEO briefs, social drafts, comparison pages, and reusable publishing workflows.
Operate and verify web apps
A workflow for agents that navigate products, fill forms, take screenshots, and verify real user flows across web applications.
Alternative shortlist
Similar skills that may fit this task.
Review a branch or diff against repository standards and the originating spec in two independent analysis passes.
Platform to build admin panels, internal tools, and dashboards. Integrates with 25+ databases and any API.
Implement work from an approved spec or ticket set, run focused and full tests, invoke code review, and commit the result to the current branch.
React and Next.js performance guidance for writing, reviewing, and refactoring production UI code.
--- name: pydantic-ai-testing description: Test PydanticAI agents using TestModel, FunctionModel, VCR cassettes, and inline snapshots. Use when writing unit tests, mocking LLM responses, or recording API interactions. ---
# Testing PydanticAI Agents
## TestModel (Deterministic Testing)
Use `TestModel` for tests without API calls:
```python import pytest from pydantic_ai import Agent from pydantic_ai.models.test import TestModel
def test_agent_basic(): agent = Agent('openai:gpt-4o')
# Override with TestModel for testing result = agent.run_sync('Hello', model=TestModel())
# TestModel generates deterministic output based on output_type assert isinstance(result.output, str) ```
## TestModel Configuration
```python from pydantic_ai.models.test import TestModel
# Custom text output model = TestModel(custom_output_text='Custom response') result = agent.run_sync('Hello', model=model) assert result.output == 'Custom response'
# Custom structured output (for output_type agents) from pydantic import BaseModel
class Response(BaseModel): message: str score: int
agent = Agent('openai:gpt-4o', output_type=Response) model = TestModel(custom_output_args={'message': 'Test', 'score': 42}) result = agent.run_sync('Hello', model=model) assert result.output.message == 'Test'
# Seed for reproducible random output model = TestModel(seed=42)
# Force tool calls model = TestModel(call_tools=['my_tool', 'another_tool']) ```
## Override Context Manager
```python from pydantic_ai import Agent from pydantic_ai.models.test import TestModel
agent = Agent('openai:gpt-4o', deps_type=MyDeps)
def test_with_override(): mock_deps = MyDeps(db=MockDB())
with agent.override(model=TestModel(), deps=mock_deps): # All runs use TestModel and mock_deps result = agent.run_sync('Hello') assert result.output ```
## FunctionModel (Custom Logic)
For complete control over model responses:
```python from pydantic_ai import Agent, ModelMessage, ModelResponse, TextPart from pydantic_ai.models.function import AgentInfo, FunctionModel
def custom_model( messages: list[ModelMessage], info: AgentInfo ) -> ModelResponse: """Custom model that inspects messages and returns response.""" # Access the last user message last_msg = messages[-1]
# Return custom response return ModelResponse(parts=[TextPart('Custom response')])
agent = Agent(FunctionModel(custom_model)) result = agent.run_sync('Hello') ```
### FunctionModel with Tool Calls
```python from pydantic_ai import ToolCallPart, ModelResponse from pydantic_ai.models.function import AgentInfo, FunctionModel
def model_with_tools( messages: list[ModelMessage], info: AgentInfo ) -> ModelResponse: # First request: call a tool if len(messages) == 1: return ModelResponse(parts=[ ToolCallPart( tool_name='get_data', args='{"id": 123}' ) ])
# After tool response: return final result return ModelResponse(parts=[TextPart('Done with tool result')])
agent = Agent(FunctionModel(model_with_tools))
@agent.tool_plain def get_data(id: int) -> str: return f"Data for {id}"
result = agent.run_sync('Get data') ```
## VCR Cassettes (Recorded API Calls)
Record and replay real LLM API interactions:
```python import pytest
@pytest.mark.vcr def test_with_recorded_response(): """Uses recorded cassette from tests/cassettes/""" agent = Agent('openai:gpt-4o') result = agent.run_sync('Hello') assert 'hello' in result.output.lower()
# To record/update cassettes: # uv run pytest --record-mode=rewrite tests/test_file.py ```
Cassette files are stored in `tests/cassettes/` as YAML.
## Inline Snapshots
Assert expected outputs with auto-updating snapshots:
```python from inline_snapshot import snapshot
def test_agent_output(): result = agent.run_sync('Hello', model=TestModel())
# First run: creates snapshot # Subsequent runs: asserts against it assert result.output == snapshot('expected output here')
# Update snapshots: # uv run pytest --inline-snapshot=fix ```
## Gates: VCR cassettes and inline snapshots
Recording or fixing rewrites files on disk. Follow this sequence; do not skip steps.
1. **Replay pass (no record/fix flags):** Run `uv run pytest` on the target path; **all green** (or failures are understood and unrelated to the artifact you will refresh). 2. **Scope locked:** Identify the cassette under `tests/cassettes/` or the `snapshot(...)` assertion to update; confirm **only** those files should change. 3. **Record or fix:** Run **one** scoped command: `uv run pytest --record-mode=rewrite …` **or** `uv run pytest --inline-snapshot=fix …` for that path only. 4. **Post-condition:** Run the same tests again **without** record/fix flags; **all green**. Inspect `git diff` — only expected `.yaml` / snapshot changes.
If step 4 fails, revert unintended diffs and fix the test or model before re-recording.
## Testing Tools
```python from pydantic_ai import Agent, RunContext from pydantic_ai.models.test import TestModel
def test_tool_is_called(): agent = Agent('openai:gpt-4o') tool_called = False
@agent.tool_plain def my_tool(x: int) -> str: nonlocal tool_called tool_called = True return f"Result: {x}"
# Force TestModel to call the tool result = agent.run_sync( 'Use my_tool', model=TestModel(call_tools=['my_tool']) )
assert tool_called ```
## Testing with Dependencies
```python from dataclasses import dataclass from unittest.mock import AsyncMock
@dataclass class Deps: api: ApiClient
def test_tool_with_deps(): # Create mock dependency mock_api = AsyncMock() mock_api.fetch.return_value = {'data': 'test'}
agent = Agent('openai:gpt-4o', deps_type=Deps)
@agent.tool async def fetch_data(ctx: RunContext[Deps]) -> dict: return await ctx.deps.api.fetch()
with agent.override( model=TestModel(call_tools=['fetch_data']), deps=Deps(api=mock_api) ): result = agent.run_sync('Fetch data')
mock_api.fetch.assert_called_once() ```
## Capture Messages
Inspect all messages in a run:
```python from pydantic_ai import Agent, capture_run_messages
agent = Agent('openai:gpt-4o')
with capture_run_messages() as messages: result = agent.run_sync('Hello', model=TestModel())
# Inspect captured messages for msg in messages: print(msg) ```
## Testing Patterns Summary
| Scenario | Approach | |----------|----------| | Unit tests without API | `TestModel()` | | Custom model logic | `FunctionModel(func)` | | Recorded real responses | `@pytest.mark.vcr` | | Assert output structure | `inline_snapshot` | | Test tools are called | `TestModel(call_tools=[...])` | | Mock dependencies | `agent.override(deps=...)` |
## pytest Configuration
Typical `pyproject.toml`:
```toml [tool.pytest.ini_options] testpaths = ["tests"] asyncio_mode = "auto" # For async tests ```
Run tests: ```bash uv run pytest tests/test_agent.py -v uv run pytest --inline-snapshot=fix # Update snapshots ```
Source provenance
Decision snapshot
recent repository activity
Audit
Install and adoption review
Agent-proven evidence
Outcome reports after resolve, review, install, and one narrow run.
No agent outcome data yet. The first agent run can report success, setup needs, risk blocks, failure, or not-relevant through /api/agent/outcome.
Install
Free and open source. Review the report before installing into production agents.
Growth loop
Scenario-led draft for pydantic-ai-testing, ready for a manual X post.
pydantic-ai-testing: Test PydanticAI agents using TestModel, FunctionModel, VCR cassettes, and inline snapshots. U... 80 stars https://www.openagentskill.com/skills/existential-birds-pydantic-ai-testing?ref=x
Listing + install path for pydantic-ai-testing: https://www.openagentskill.com/skills/existential-birds-pydantic-ai-testing?ref=x Install: npx skills add existential-birds/beagle --skill pydantic-ai-testing
Listing source
This listing was indexed from public sources and is not marked official until a maintainer claim is approved.
Attribution links to the public repository or creator profile. Creators can claim the listing to update ownership signals.
Claim this skillOwner claim
This Registry indexed listing is attributed to existential-birds but is not marked official yet. Claim it to add a verified owner signal and make future launch, install, and audit updates easier to trust.
Creator backlink kit
Show the canonical listing, current trust and audit signals, and real Agent-Proven evidence where developers evaluate the repository.
[](https://www.openagentskill.com/skills/existential-birds-pydantic-ai-testing?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)
[](https://www.openagentskill.com/skills/existential-birds-pydantic-ai-testing?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)
[](https://www.openagentskill.com/skills/existential-birds-pydantic-ai-testing/audit)
[](https://www.openagentskill.com/skills/existential-birds-pydantic-ai-testing?ref=github&utm_source=github&utm_medium=referral&utm_campaign=creator_badge)existential-birds
@existential-birds
Share whether this skill looks useful for your agent workflow. Aggregated feedback improves rankings over time.
Sandbox only
Code Review
Review a branch or diff against repository standards and the originating spec in two independent analysis passes.
168.6K StarsAppsmith
Platform to build admin panels, internal tools, and dashboards. Integrates with 25+ databases and any API.
40.8K StarsImplement
Implement work from an approved spec or ticket set, run focused and full tests, invoke code review, and commit the result to the current branch.
175.7K StarsVercel React Best Practices
React and Next.js performance guidance for writing, reviewing, and refactoring production UI code.
30.9K StarsPermission surface
shell or command execution, network or browser access
Agent outcomes
No agent outcome data yet
Docs
Usable metadata, review docs
Risk summary
Install readiness