Lint agent skills so their code examples are real, runnable, and schema-accurate — including anti-gaming rules

Agent skills guide coding agents through complex APIs via code examples. Generate many skills and the agent hallucinates examples, destroying accuracy. Lint the skills themselves: extract code blocks, type-check them, run them in a sandbox, verify against the OpenAPI schema — then keep tightening the rules in a whack-a-mole loop as the agent games each one.

From AI Engineer — Stop Prompting — Greg Pstrucha, Sentry

Problem: Hallucinated code examples in generated skills tank agent accuracy; skills drift when the underlying API changes.

For: Teams building agentic products that rely on generated skills/documentation with embedded code examples.

Products from this video

ast-grep Clippy ESLint Flake8 Seer Sentry Taskless

Examples

🔒 Full analysis locked

Unlock more videos and the full analysis

A credit unlocks one video's full analysis for good — the build steps, the tools and how each was used, the methods behind every use case. Pro opens the whole library instead, and raises how many videos you can analyse a day.

Unlock full analysis — free

Other takes on AI evaluation and benchmarking

All AI evaluation and benchmarking ideas →