AI product Open source

OpenAI Evals

OpenAI Evals is an open-source framework for evaluating large language models and systems built with them, including tool-using agents and prompt chains. It includes a registry of benchmark evals, supports custom model-graded and private evals based on a user's data, and uses a completion-function protocol for advanced workflows. The package can be installed with pip and run locally with an OpenAI API key; the repository also documents optional result logging to Snowflake and configuration through the OpenAI Dashboard.

View repository Mentioned in 1 video ↓

What OpenAI Evals is used for

1 use taken from transcripts — each links to the moment in the video.

  • Evaluations measure whether agents produce business results such as conversion, customer value, satisfaction and re-engagement.

Videos mentioning OpenAI Evals

1 in the library.