AI product
Assistant Benchmark is a comparative evaluation of AI assistants created by David Pawlan. It gives different assistants the same real-world prompts, including tasks such as email management, travel booking, and financial administration, and compares their outcomes, follow-up questions, speed, and performance across 16 dimensions. The benchmark is based on ongoing testing of multiple assistants and examines both task execution and the degree of autonomy assistants can usefully take.
1 use taken from transcripts — each links to the moment in the video.
Compared AI assistants by giving them the same real-world prompts and evaluating outcomes, follow-up questions, speed, and performance across 16 dimensions.
1 in the library.