AI product
An evaluation toolkit for Strands AI agents that measures off-script behavior, violations of behavioral guardrails, and other incorrect behavior against test and live data. It is associated with Strands Agents, an open-source AI agent SDK for Python and TypeScript.
1 use taken from transcripts — each links to the moment in the video.
Evaluates how often agents go off script, violate behavioral guardrails, or produce incorrect behavior against test and live data.
1 in the library.