← All transcripts

Your LLM Judge Is a Confident Liar: Building Better Verifiers — Browserbase Transcript, AI Summary & Key Points

AI Engineer · 3 days ago · Science & Technology · 21:14 · EN

Watch on YouTube

AI Summary

LLM judges for web-agent benchmarks can overestimate success because they use weaker models, lack task-specific rubrics, inspect irrelevant or too many screenshots, and may omit final answers or action histories. The Universal Verifier generates criteria for the requested task, ranks relevant screenshots as evidence, isolates rubric errors, checks for hallucinations, and separates controllable from uncontrollable failures. On the same Fara 7B model and benchmark, the official WebVoyager judge gave a 74% success rate while the Universal Verifier gave 38%. The Universal Verifier reduced false positives from almost half to zero and reached human-level agreement. Higher-quality verification also produced better filtered training data and higher-quality models. An autoresearch loop built a verifier in about one day instead of three weeks but reached about 70% of the agreement level, showing that human input still improved the result.

Key Points

  • Browserbase provides cloud browser infrastructure and runtime for deploying web agents at scale, while Microsoft Research collaborated on building verifiers.
  • Deterministic web-agent evaluations stopped scaling because the web is open-ended, has multiple valid paths, lacks consistent ground truth, changes over time, and introduces environment errors.
  • LLM judges bundled with benchmarks such as OSWorld, Online-Mind2Web, and WebVoyager can be confidently wrong, so optimizing against them can produce a more competent liar rather than a better agent.
  • Fara 7B received a 74% success rate from the official WebVoyager GPT-4o judge, but the Universal Verifier scored the same model and benchmark at 38%.
  • Existing verifiers use smaller models, often lack rubrics, may inspect irrelevant screenshots or overflow their context window, and sometimes omit the final answer or action history.
  • The Universal Verifier generates a rubric with criteria for the task, ranks the most relevant screenshots for each criterion, checks those screenshots against the agent's claims and environment state, outputs partial process scores, and produces an outcome true-or-false value.
  • Rubric creation — Grade only what the task asks for and do not add extraneous criteria that artificially lower scores.
  • Error isolation — Do not cascade an error from one rubric criterion into later criteria; an incorrect choice of person should not invalidate an otherwise accurate answer about that person's net worth.

AI in practice

Agents

  • Build a verifier that agrees with human labels. 2 held 14:46

Tools & resources

1 item

Links mentioned

🔒 Full analysis locked

Unlock more videos and the full analysis

A credit unlocks one video's full analysis for good — the build steps, the tools and how each was used, the methods behind every use case. Pro opens the whole library instead, and raises how many videos you can analyse a day.

Unlock full analysis — free

Transcript

Searchable transcript of Your LLM Judge Is a Confident Liar: Building Better Verifiers — Browserbase — AI Engineer (21:14). Search for a phrase, then click its timestamp to jump straight to that moment in the video.

Captions sourced from the original video on YouTube, published by AI Engineer. The video, its captions and all related intellectual property remain the property of their respective owners; AINotes claims no ownership. Provided for research, accessibility and search — see the Transcript Notice and Copyright Policy.

00:01 [music] Good afternoon everyone. Um, thank you for coming. Hope you've had a great expo so far. We're in the last stretch, but I know it's been a wonderful, wonderful expo, at least for me. So hopefully you're learning a lot and we can teach you a little bit more. Um, so I'm Miguel. I'm the tech lead for uh our agent platform at Browserbase. For those of you who don't know, browserbase is a infrastructure company for deploying agents on the web.

00:37 Um it basically unlocks all of the web, not just what is easily accessible uh by managing browsers in the cloud and all the runtime that uh requires to do agentic web automation at scale. So um let >> and I'm Corby. I'm a researcher at Microsoft and I've been collaborating with with browser base to help build really good verifiers. >> So today we're here to talk about some research that we published about a month ago.

01:06 Um and it is regarding evalance. It is regarding RL environments and LM judges as verifiers. So early on so I've been working on this open source framework called stchen for about a year and a half now. And from the beginning as any good agentic product needs to have we were building evals to make sure that the changes that we were introducing in the framework and that the model capabilities continue to improve and hill climb on those evals and at the beginning all those evals were deterministic environments.

01:41 um it was relatively simple. Model capabilities were some somewhat constrained and so we were able to build a lot of deterministic environments as model capabilities continue to improve as we continue to make improvements on the agent harness that just didn't scale. We found ourselves constantly patching or generating static sites um coming up with longer trajectories with a lot of checkpoints to make sure that the evals were deterministically verifiable.

02:09 And so there's many different project problems with automating the web at scale and verifying that what you intended to do actually succeeded. And it's that the web is very open-ended. There's not just one path to correctness. Um there is no ground truth. Sometimes the web changes and the product that you were checking for before is no longer there.

02:31 So that breaks the whole determinism of your environment and there's blockages. there's a lot of the environment errors that you face that don't let you get sort of a verifiable deterministic signal. So we recurred to what most people do which is trying to use an LLM judge to help improve um and scale this operation. But we quickly found out that a lot of the leading benchmarks and some of you may know about OS World uh for computer use but there's some analogous benchmarks for web use called online minor web and web

03:07 voyager. Those come with LLM judges as verifiers. And what we noticed when we were using our sort of human expert annotators to verify the verifier is that in many instances they're very confidently wrong. And that leads to results that you can't really trust. So very quickly we realize that even if you were using this as an RL reward or as a signal to improve your hardness and auto research, you're not really training a better agent.

03:40 You're just training a more competent liar. And so in a real world use case at Microsoft, we trained this FARS 7B model which is a a web browser agent. And the same model on the same benchmark, we judged it according to the official web voyager judge, which is GPT40. Uh and it said that it had 74% uh success rate. But there's a huge gap between the real truth which is when we used our new universal verifier uh which has high agreement with human labels that that number quickly becomes 30 like 38%.

04:14 So there's a huge gap between what the existing verifiers say and what is actually the ground truth. And so we needed to find a way to to close this gap. The existing verifiers have lots of weaknesses. Um they're they use smaller and like much dumber models. uh as the LM as a judge they use 04 mini or GBT40 they don't even use rubrics so rubrics are the most critical thing that you need to have because you need to assign credit to where credit is due they don't look at the the relevant screenshots or they or they try

04:46 to look at all the screenshots and they quickly get lost and overflow the context window of the LLM as a judge and sometimes they don't even look at the final answer or the action history of the model so our verifier checks all of these boxes um and this is kind of how it works uh at a very high level you And you can look in the paper for more details, but given a task like book a cheapest flight from Seattle to Boston, the first thing we do is we generate a really good rubric and I'll give some examples of that.

05:11 And a rubric has like maybe 10 different criteria of what success looks like. And then given that that criteria in the rubric, we look at the agents trajectory. All of the screenshots in this case are the the evidence of the ground truth. and we rank what are the most relevant screenshots in the trajectory for each criterion. And then we use that group of of top K evidence to determine whether that criterion was met or not and if there's any like differences or contradictions between what the agent said it did and what

05:43 the state of the environment actually showed. And then given that we output uh a scored rubric which we we call like a process score because it gives partial credit in some cases. and we give an we output a uh an outcome boolean uh value to turn to basically say whether the agent accomplished the task according to what a reasonable user was would expect.

06:08 Um and so this is an example of what the rubric uh the rubric output is. It's basically a list of criteria um as well as the outcome verifier is basically a true false flag with an explanation as to why a reasonable user would expect um this trajectory to have succeeded or not. There were four guiding principles when we created this universal verifier.

06:33 One is around rubric creation which is we wanted to grade only what was asked and not any extraneous criterion. Uh we also did not want errors to cascade from one rubric uh criterion to another. We wanted them to be isolated. And I'll give an example of that. You also wanted to look at the ground truth screenshots. Like this is a must. The agents will often overconfidently claim that they did something when in fact they did not do it.

06:56 And so you have to look at the ground truth state. And then we also needed to separate um what the agent could control versus what the agent could not control. The agent operates in a web browser. The web browser is an environment that it doesn't always have ability to control and I'll show some examples of that too. So this is an example of what a good and bad rubric looks like.

07:15 The task here is to uh find a cheap uh hotel in Jakarta for these dates and then use the hotel's address to search for the closest coffee shop and then output the name and address of that coffee shop. Now a bad rubric which we did see happen would output a criteria saying like please tell me the total price for the stay at the hotel. This is actually an extraneous criterion that was not asked for in the task but we saw a lot of rubrics naively generated do something like this and they would artificially deflate the

07:46 scores because the agent didn't do something it wasn't asked to do. Um when it comes to not cascading errors, um there are very subtle mistakes because a lot of tasks build on each other uh if they are like multi-step tasks. So in this example here, the task was to um determine the net worth of the individual with the longest last name from among the members of NSync and Backstreet Boys.

08:12 Okay. And so the agent in this case mistakenly thought that Timberlake had the longest last name which is only 10 letters whereas Kirkpatrick actually had the longest last name. So the model made one arithmetic mistake in one criteria of the rubric and that should not cascade to the next criteria of the rubric which is report the net worth of of the person that you chose.

08:34 Right? So if you're able to report the net worth of Timberlake uh accurately then you shouldn't get penalized in this criteria. you should only get penalized in that one. Hallucinations is the biggest one uh that we really wanted to focus on. So in this task, this is an example of a very subtle hallucination uh where you wanted to find some information about some image captioning model and the agent said that this model had plus 6.2% in some cider score, but in reality the abstract of the underlying paper said it was

09:06 only 2.8% in cider score. And so this was a like a very subtle mistake that even humans will not catch unless an LLM will flag it. Um so this kind of gives uh an overview of of how process scores and outcome scores differ. Um in the case we mentioned one of the principles was we wanted to separate controllable versus uncontrollable um failures. So, if the task is to buy, say, this uh plushy toy from Amazon, and the agent accurately searched for the plushy toy, but found that it was out of stock, it could not continue.

09:39 And so, here we say that the agent gets full marks for doing its best effort uh to achieve the goal, but the outcome was still not met because it couldn't buy the thing that it wanted to buy because it was out of stock. Um, so we basically enumerated a bunch of scenarios as to what could happen and how do we assign credit if the agent was able to find a similar alternative plushy toy in a different way.

10:04 Um, it would still get success there uh for both for both cases and it would get penalized if it made a controllable mistake. Uh, it would get penalized if it made hallucinations and so on and so forth. So we basically enumerated a a schema of like all the possible failure modes and how we would assign credit to that and the universal verifier adheres to those to those schema.

10:25 So >> basically in in order to evaluate the verifier and really tell whether or not we were able to hill climb the accuracy of that verifier, we started designing a lot of experiments with human experts and built a whole platform to collect data on what we call the ground labels that we've also released under cool verify bench for others to train upon.

10:49 But the process that we followed was the following. At first, we would show the full trajectory of evidence to the human verifiers and the annotators would judge based on the evidence and the result of all the actions that the agent took. After that, they would be prompted to a seeing the judgment of the universal verifier to say whether they agree or disagree with a human with a universal verifier.

11:18 And we found this to be a very reliable signal to tell whether or not we were working in the right direction because I'll jump to this one straight up, but the verifier started correcting humans at one point. They it would identify things that the humans were missing. And so ultimately the result of the of the universal verifier on these golden labels was that we were able to reduce false positives from almost half to zero and coins cappa isometric to measure inner annotator inner annotator agreement.

11:56 Um and what we can see here is that the h universal verifier agrees with humans as often as humans agree with one another. So we can see that 0.58 coins scalpa score uh versus the two annotators per task um where we computed coins scappa as well to see how much they agreed with one another. >> Another way another way to uh verify the another way to verify the quality of our verifier is to actually train on data filtered by it.

12:31 So what we did was we did ran an experiment where we held constant the number of training examples. In this case it was 3k or 9k training trajectories, but we filtered those training trajectories based on whether our verifier said they were pass or not. And so when you're doing SFT, which is what we did here, you want to train only on SFT examples where they were true uh like successful trajectories.

12:52 And so if you train on 3,000 trajectories filtered by the process score of our universal verifier, you will get a much higher quality model than if you were to train on a worse verifier like our own a baseline verifier. And this held at larger scales as well. And so this basically this experiment told us that uh if you have a higher quality verifier, it means you can filter higher quality data uh and training on that data leads to a higher quality model.

13:22 And so this was like the ultimate experimental proof in addition to the um uh to the to the human agreements uh that this model this verifier is working and is and is reliable um in a real production uh scenario. Um the last thing we did kind of as a side project is we wanted to determine whether AI in auto research can build the same verifier that we built.

13:46 So basically what I did and what Miguel and I did is we sat down for three weeks um to build the universal verifier and to tweak the prompts and to write the code for it. And we basically ran about 30 different experiments over the course of three weeks to determine whether the verifier we were building agrees with human labels. Okay. And so that's what the Coen Kappa score here on the Y-axis is measuring.

14:08 It measures whether an individual verifier system is agreeing with the human labels. And so the blue line here is the model Miguel and I, the system Miguel and I were were creating this auto research uh this uh universal verifier. And then we wanted to see if if we stripped everything that we did away, could AI um in an auto research loop build the same verifier that we built with the same level of quality and fidelity?

14:36 And that's what the red and green lines here are. The red line is if we took away all of our code and prompts, could AI replicate that? The conclusion that we got from this is that while we took about three weeks to build this verifier, uh an auto research loop could do it in about one day. It ran the same number of experiments in about one day. However, it only reached about 70% of the agreements that our verifier uh was able to reach.

15:01 So there's still a gap there. And so the kind of the conclusion that I would draw here is that uh you can use auto research to build uh metrics and verifiers um at least to help you speed up uh experimentation. Um but you probably still need some level of human intervention here. Um then again, these these results were from Opus 4.6. I haven't tried it with Fable yet.

15:24 Maybe it will do better, but it's a pretty powerful baseline and it got pretty far um along the way. So, not only can auto research help you build a model, but in this case, auto research can help you build a verifier as well, which we thought was a pretty cool um thing. And and building verifiers is as important as building the models themselves is one of the takeaways I want to leave you with.

15:45 So >> yeah, and just to double emphasize um the green light is the auto research verifier um primed with a lot of the findings that we had gotten from the three weeks of experimentation and in that optimization it was able to reach higher percentages than we were ever before. So combining that human intuition with auto research loops has proven to be the most powerful recipe for many research tasks.

16:12 Um so the all of the work is open source. The paper is published under this uh as a preprint um on archive. Um and the Microsoft repo contains the golden labels on kua verifier bench. It contains the code to run the experiments. But beyond that um this research also yielded an entire new benchmark. So we've talked briefly about online minor web and web voyager.

16:46 I spend a lot of time working with labs helping make sure that we can provide them quantitative signal to improve their models. A lot of those open source benchmarks are very quickly getting saturated and it is very difficult to provide any sort of signal because it feels like the the test data is already in the distribution of the training data. And so this new benchmark has proven to have the biggest gap to complete success and it's my daily driver to make sure that I can tell whether model capabilities are strong,

17:22 whether they lack and how to continue to close the gap. So >> did you have more slides? >> No. >> Okay, >> we can take a little bit of I think questions, but uh we're good on time. >> Yes. Yeah. So the question is the question is whether we've used the uh the universal verifier for other kinds of tasks beyond just web web tasks. So one thing we're working on in Microsoft right now is to apply a version of this verifier for desktop tasks like things that you would do on your laptop like kind of enterprise workflows.

18:00 And so we hope to publish another benchmark uh of not just web data but also like enterprise style desktop data. Uh and we would have basically the similar the same kind of verifier for that too. It would look at a little bit more information because on desktops you have the terminal, you have more telemetry, more logs than just the browser. So the verifier would take those things into account as well.

18:22 Yeah. Any other questions? >> Yeah. >> So I personally never when you present I was thinking about like how do you think about that show? Yeah. >> Yeah. So, the question is about how do you make sure you're not overfitting the verifier to the human labels because that is a problem. Uh Miguel and I spent a lot of time uh working on this question. And so like one thing that we did was we uh kind of held out uh two sets of labels.

19:19 One we we basically had a set of like 150 trajectories. um on about 50 of those trajectories, I labeled them myself and I used them to hill climb in those 30 experiments. But for the other hundred trajectories, they were basically held out to us. They were done by humans uh that we had paid uh with like 2x overlap. So once our verifier was done being built, we gave it to those human annotators and each human labeled each trajectory, sorry, two humans labeled each trajectory and then that's what we published in KUA

19:52 verifier bench. And so those are the numbers that we report here. Um so we're pretty confident that we didn't overfit because we held out at least twothirds of the labels. Um like we did not train like iterate on them. >> Yeah. So that's very important is you don't want to train on like to be able to >> Yeah. Yeah. Yeah. The gold standard is uh is to hire humans um and train them to do the the verification task themsel and then see whether your system agrees with it.

20:34 Yeah. Any other questions? Anyone want to help us build more benchmarks? >> Yeah. Oh yeah. Okay. Okay. Yeah. Yeah. Yeah. >> Awesome. Well, thanks a lot. Uh hopefully you enjoyed the rest of the expo. It was a pleasure. >> Thank you so much. >> Thank you so much. >> Yeah. Yeah.