AI product Open source
OpenAI Evals is an open-source framework for evaluating large language models and systems built with them, including tool-using agents and prompt chains. It includes a registry of benchmark evals, supports custom model-graded and private evals based on a user's data, and uses a completion-function protocol for advanced workflows. The package can be installed with pip and run locally with an OpenAI API key; the repository also documents optional result logging to Snowflake and configuration through the OpenAI Dashboard.
1 use taken from transcripts — each links to the moment in the video.
Evaluations measure whether agents produce business results such as conversion, customer value, satisfaction and re-engagement.
1 in the library.