AI did not make shipping faster when human review and verification remain the bottleneck; exhaustive verification can let teams ship at the speed their AI agents write code.
Searchable transcript of Why AI Didn't Actually Make You Ship Faster — Gabriel Spencer-Harper, Meticulous — AI Engineer (11:15). Search for a phrase, then click its timestamp to jump straight to that moment in the video.
Captions sourced from the original video on YouTube, published by AI Engineer. The video, its captions and all related intellectual property remain the property of their respective owners; AINotes claims no ownership. Provided for research, accessibility and search — see the Transcript Notice and Copyright Policy.
00:01 [music] Hello everyone. Hi. Um, my name is Gabe. I'm one of the co-founders and CEO of Meticulus. Um, I'll just wait one more minute for people to come in and then we can and then we can get started. Um so in today's demo or talk I'll talk about verification of code. Um I'll spend one minute introducing meticulous and then one minute giving some background context on the company and then sort of eight minutes giving a demo and technical technical overview.
00:40 Um so if you have exhaustive verification then you can ship code at the speed that your agents write it. If you don't, then someone somewhere at your organization is spending time verifying code and that is now your new bottleneck and you're forced into a position where you're trading off velocity against bugs and user experience. So I'll give um super quick context on Meticulus.
01:15 So, Meticulus gives you exhaustive or near exhaustive verification of your front-end codebase with zero developer effort. Um, it's deployed across the entire engine organization at all of these companies like Discord, Whiz, Dropbox, Notion, 11 Labs, Launch Darkly, and many, many other great companies. So, if an engineer touches front-end code at one of these companies, they use Meticulus every every day.
01:42 However, before we get into the product and talk more about verification, um, let's just cover the current sort of state of the world. So, AI writes code faster than humans can review it. And review and verification is now the new bottleneck. The other state of the world is that assertionbased testing alone isn't enough. So no matter how hard a human or an agent tries to define the correct behavior of software a priority upfront in assertions the space of possible regressions is too vast to exhaustively cover.
02:28 And so a good clarifying question is if an engineer in your organization uh AI generated this pull request, would you feel comfortable hitting the the merge merge and ship button? And in most organizations, the answer is no. You would want to go and do the hard work of checking out all the different feature flags, all the different roles, permissions, settings, configurations, and the various different edge cases in order to understand the exhaustive impact of this of this change and see if you can spot any issues.
03:04 So there are three sort of like downstream consequences of that current state. The first one is bugs or regressions uh which have a a business impact. The second one is engine organizations will spend doubledigit percentages of their time maintaining end to-end test suites. So across manual validation, review, uh debugging flakes or updating or maintaining the test suite, um all spend a lot of time maintaining maintaining maintaining tests.
03:36 And then the third like problem or downstream downstream consequence of of the state of the world is that um if you if you can get exhaustive verification then you can program in a new and different way. So all of a sudden you can bump all your dependencies, you can do sweeping refactors, you can make any AI generated change with complete complete confidence.
04:03 So then the question becomes um that sounds great. How do you how do you get exhaustive verification? What does it look like and how do you how do you get there? Um so I'll now give you a demo of how how meticulous meticulous works. So, we give you uh give me one second to set this up. Great. Um, so we give you one line of JavaScript to inject onto nonproduction environments.
04:41 So that's localhost, QA, dev, staging. That JavaScript instruments the web browser and records thousands or tens of thousands of flows or workflows. So someone clicking the login button, clicking the settings panel, clicking the analytics panel. And then when you open a pull request on your CI runner, you spin up your web application on localhost 3000 and Meticulus takes a subset of those recorded workflows and it replays them.
05:10 So it dispatches each event one by one. And as it dispatches those events, so click the login button, click the senses panel, click the analytics panel, it takes a screenshot at each atomic moment throughout time. And so it generates a sequence of screenshots for every workflow, one before your code change and one after your code change. It then diffs those two sequences together to show you what would be different about your application if you were to merge that code in before before you do.
05:38 So in this example, there's two changes here. One is someone's changed the name and the second is they've introduced an error to a to a drop down. And so when you open a pull request, Meticulus posts a comment typically within a few minutes. And when the developer clicks into this comment, the Wi-Fi is slow, so I'm just going to switch tabs. But when a developer clicks into this comment, Meticulus will not tell you whether or not there's a bug.
06:09 What it will show you is a set of diffs that show you what's different about the application if you were to merge that code in. And so each diff has a before state and an after state. And so by reviewing the diff, you can determine whether that change is expected or unexpected. So here's the name change on light mode. Here's the name change on dark mode.
06:31 And then this is a manifestation of that logical logical error. So typically when you write an assertionbased test, you um use your business context and judgment to encapsulate the correct behavior of software up front and and encapsulate that into your assertion based test. In meticulous, we just show you what's different and you exercise your judgment at review time or or or an agent agent does.
07:00 Um give me one second to switch this back. Cool. Um, so there are sort of three technologies that are really important for helping build a mental model of the tool and how it works, how it works and why it works. So the first one is that everything is mocked out. So um at record time we record all the network requests and responses and then at replay time we stub those in and we perform that mocking for two reasons.
07:37 One is to make the test item potent. So you can run them a thousand times in a row and get the same results each and every time. And the second reason is it means that every test is isolated from every other test which eliminates the risk of race conditions and allows you to paralyze the test horizontally and get get the results in it in a few minutes.
07:55 The second key technology is that you get radically less or orders of magnitude less flakes with meticulous than any other tool including cypus or playright. And the reason why is we augment the browser from the scheduling engine layer up to be fully deterministic. So when browsers were first designed determinism was not a design goal. And so there's many different sources of randomness.
08:19 For instance, an animation spinner depends upon the clock speed of the CPU of the machine rendering the browser or the interle between set timeout and set interval. And so we handle all of that randomness and uh make the browser deterministic or near deterministic so that you have radically radically less flakes. And then the third technology which is the most um important one for grocking like why this works is if you want this to be exhaustive it has to cover every feature plan combination every permission every role
08:54 every config every possible branching through the application. So how do you how do you do that? So one technique that we use is that for every session or flow that we record, we replay it once against main branch or master branch and we monitor what lines of code get executed and then we build a map or an index that maps each individual workflow to lines lines of code executed and choose a subset of workflows that maximizes code coverage across your application.
09:21 And typically uh give me one second. Typically I would say that code coverage in our opinionated view is a terrible metric for every tool in the world apart from meticulous and the reason why is you could write a cypress test or playright test that covers 100% of your codebase but only makes a single assertion and so code covered is not the same as code tested but with meticulous it actually is approximately the same because we do this really unusual and radical thing which is we're taking a screenshot at every
09:57 possible moment throughout a flow. So if there's a 412 set flow and there's a single pixel diff, meticulous meticulous will will will flag it. And so you can start to see how all these technologies layer together. Um you only get exhausted verification because of this code coverage algorithm. That algorithm only really works because you take a screenshot at every possible moment.
10:20 If you take a screenshot every possible moment, you're taking on the order of tens, hundreds of millions of screenshots. And so, traditionally, you would just drown in noise or drown in flakiness. And so, you have to solve flakiness at the root level. Um, which is which is what we what we what we do by by augmenting augmenting the browser. Um, these are some nice things that our customers customers have said about us.
10:45 Um, and that concludes today's talk. If you're interested in finding out more then come to booth LG6 and thank you so much for listening. [applause] >> [music]