← All transcripts

How VS Code Went from Monthly to Weekly Releases with AI — Harald Kirschner Transcript, AI Summary & Key Points

AI Engineer · 6 days ago · Science & Technology · 19:35 · EN

Watch on YouTube

Answer

VS Code moved from monthly to weekly releases by rebuilding its delivery system around AI: making the codebase agent-ready, speeding up CI, giving agents access to the application, enforcing AI code review, automating issue and error triage, staging rollouts, and learning through evaluations and rapid prototypes.

AI Summary

VS Code moved from a monthly to a weekly release cycle after more than 10 years of monthly releases by evolving its entire software delivery system around AI. The approach combines agent-ready codebases, reusable skills, faster TypeScript Go builds, application-level feedback loops, mandatory AI code review, AI-assisted issue triage, telemetry-driven fix PRs, staged rollouts, VSC-Bench evaluations, daily prototyping, and smaller squads. Agent-written code survival increased from 55% with GPT-4.1 to 86% with Claude Opus 4.6. The central lesson is to improve the feedback and quality systems around agents rather than simply use more AI.

Key Points

  • VS Code moved from monthly releases to weekly releases after more than 10 years of monthly releases.
  • Agent-written code survival rose from 55% with GPT-4.1 to 86% with Claude Opus 4.6, reflecting increased trust in AI-generated code.
  • AI increased both useful community contributions and low-quality issues and pull requests, creating a need for stronger quality and triage systems.
  • Shipping faster, shipping quality, and learning faster require evolving the complete software delivery system rather than simply using more AI.
  • Agent-ready codebases use lightweight agents.mmd documents, repository maps, documentation, onboarding material, and living instructions that evolve as agents make mistakes.
  • Accessibility expertise was encoded into a maintained skill so the whole team could apply practices that previously depended on one specialist.
  • A practical test for an agent-ready VS Code repository is whether a product manager can effectively vibe-code in it and produce work that engineering can accept.
  • TypeScript Go improved builds by 10 times; slow CI and other bottlenecks compound when 10 or 20 agents hit them simultaneously.

Tools & resources

7 items

ENo. 0707
AIAINotes.us Tool

Electron

Open source · electron/electron

Electron is an open-source framework for building cross-platform desktop applications with JavaScript, HTML, and CSS. It is built on Node.js and Chromium, provides binaries for macOS, Windows, and Linux, and is used by applications such as Visual Studio Code. The project is MIT-licensed, maintained under the OpenJS Foundation, and distributed via npm.

Mentioned in
5 videos
Kind
Other
GNo. 0495
AIAINotes.us AI product

GitHub Copilot

github.com/features/copilot

GitHub Copilot is an AI-powered software-development tool created by GitHub, owned by Microsoft, in collaboration with OpenAI. It suggests code completions, whole-line and whole-function snippets, documentation, and software changes for review by an experienced engineer in supported editors and IDEs. Its workflow includes inspecting changes line by line, assessing architectural correctness, and coordinating work with other pull requests. It is offered as a commercial subscription integrated with Visual Studio Code, Visual Studio, JetBrains IDEs, and GitHub Codespaces.

Mentioned in
11 videos
Kind
AI
PNo. 0027
AIAINotes.us Tool

Playwright

Open source · microsoft/playwright

Playwright is an open-source framework for web testing and browser automation, developed in the Microsoft repository. It drives Chromium, Firefox, and WebKit through a single API and includes an end-to-end test runner with isolated browser contexts, auto-waiting, web-first assertions, resilient locators, parallel execution, and reusable authentication state. Its tooling also includes a browser-automation library, CLI for coding agents, MCP server for AI-agent and LLM-driven automation, and tracing that records actions, DOM snapshots, network requests, console messages, screenshots, and videos for debugging.

Mentioned in
8 videos
Kind
Other
TNo. 4874
AIAINotes.us Tool

TypeScript Go

In the AINotes directory

TypeScript Go is a TypeScript build tool described as speeding up builds and reducing the CI/CD bottleneck that occurs when many coding agents run concurrently. The video attributes a roughly tenfold build-speed improvement to it.

Mentioned in
1 video
Kind
Other
VNo. 0374
AIAINotes.us Tool

Visual Studio Code

Open source · microsoft/vscode

Visual Studio Code is a free, open-source Microsoft-developed source-code editor and development environment for Windows, macOS, Linux, and the web, built from the Code - OSS repository. It supports code editing, navigation, code understanding, lightweight debugging, integration with existing development tools, and extensibility through extensions. It provides an environment for building with AI agents that plan, modify, and debug code, managing multi-agent workflows, hosting AI coding plugins such as Claude Code, and running workflows such as Spec Kit commands.

Mentioned in
11 videos
Kind
Other
VNo. 4876
AIAINotes.us AI product

VSC-Bench

In the AINotes directory

VSC-Bench is an evaluation tool for testing agentic scenarios in Visual Studio Code. It runs product evaluations and supports adding new scenarios from GitHub issues.

Mentioned in
1 video
Kind
AI
XNo. 4875
AIAINotes.us AI product

Xcode MCP

In the AINotes directory

Xcode MCP is an MCP tool that lets an AI agent create an iOS app in VS Code, open and interact with the app, and capture screenshots for a feedback loop.

Mentioned in
1 video
Kind
AI

AI in practice

Used for

Agents

  • Triage incoming GitHub issues by filtering spam, enriching and translating reports, and assigning them to area owners. 2 held 12:36
  • Turn production error stacks into diagnosed issues and pull requests that attempt to fix the underlying problems. 2 held 13:23
  • Bootstrap a new VSC-Bench evaluation scenario from a GitHub issue. 2 held 16:24

Links mentioned

🔒 Full analysis locked

Unlock more videos and the full analysis

A credit unlocks one video's full analysis for good — the build steps, the tools and how each was used, the methods behind every use case. Pro opens the whole library instead, and raises how many videos you can analyse a day.

Unlock full analysis — free

Transcript

Searchable transcript of How VS Code Went from Monthly to Weekly Releases with AI — Harald Kirschner — AI Engineer (19:35). Search for a phrase, then click its timestamp to jump straight to that moment in the video.

Captions sourced from the original video on YouTube, published by AI Engineer. The video, its captions and all related intellectual property remain the property of their respective owners; AINotes claims no ownership. Provided for research, accessibility and search — see the Transcript Notice and Copyright Policy.

00:12 Hey uh everyone. This is a hard tone to follow. Why did everybody leave? No. Um this is the rerun of a talk that she gave yesterday upstairs. So this will be better than the one yesterday. So uh kudos to you for catching this one. I'm working on VS Code and we move from a monthly release to a weekly release cycle thanks to AI. And this is about the things we learned, the systems that made this work, not just to ship faster, but actually to ship more effectively.

00:43 So, it's not about how we build agents into VS Code, which we can try out anytime. You can come by the booth and see us showing of agents, but this is about how we use agents to build the product itself. And that's VS Code, that's the copci and other areas, but also how as you use agents and more AI, things are breaking and you have to rethink actually how you work.

01:05 And that's a very small team shipping to over 50 million users. What happened is over the past year is we started tracking already early on this code survival metrics in VS code. So that means the percent of agent written code that actually gets committed. how much kind of garbage has the agent produced that you threw out the window again or didn't trust and that you that you read it u and and scrapped.

01:32 So GBD 4.1 started at 55%. But then over time and over improving the harness and new models cloud opus 4.6 is now today at 86%. And that increase clearly shows developer trust and confidence in shipping AI code faster and more AI code generated. And this success created problems on the VS code side. We're one of the biggest or the biggest open-source project on GitHub.

02:03 So we saw a massive increase in issues because AI helps to file issues. So we actually see an increase in quality issues but also the counterpart of more automated lowquality issues. We also see more open PRs by the team by our own velocity and by our community. And interesting part we actually see everybody would assume we see a lot of garbage PRs but actually the number of community merged PRs is going up as well which is really amazing that we can unlock more community contributions as well.

02:36 And as part of this new workflow, we actually moved several months ago, we moved from monthly releases, which happened over 10 years. Since the inception of VS Code 1.0, we have been shipping a new version of VS Code on a monthly basis. Including fun facts like the January release actually ships in February. So when you read the release notes in VS Code, you're wondering why it's February and you're reading the January release notes if you're a month late.

03:02 No, that's just how we tracked our iterations. So that cadence was to match to velocity, but it also showed us that the it's not about using more AI as you work dayto-day, which a lot of companies are pushing. Like you're not token maxing. You're not spending every day using AI. It's about evolving the whole system, how you're shipping software to evolve to actually make better use of AI along the whole process.

03:33 And that what makes this velocity actually work and doesn't just generate more code. I'm breaking this down into this first part which a lot of people have been experiencing. I see shipping code faster and that's the easy part. You have like this 100x engineers that has built this beautiful skill MCP plug-in system and they're just shipping code like crazy and slowly maybe more in the more people in the team uh adopt the same patterns and you see a lot more code shipping faster.

04:03 But the problem is once you ship faster, you get into code review. You get into how do you actually ship high quality code that doesn't constantly break. And that's important for us as VS Code. We're shipping binaries to user machines. So if something breaks, that recovery is even more costly. But once you hit that velocity with shipping high quality code faster with confidence, how do you actually know you're shipping the right thing or if we should ship this thing?

04:32 And that's where AI kicks in to actually learn faster in the end and shipping better products which a lot of people are missing out as as they write more code is applying that product taste and applying that learning. So let's start shipping faster. Now that everybody can vibe, we can really all ship faster. And it's a big part is getting your codebasation ready.

04:56 So that's kind of the no-brainer slide that everybody should already been doing is thinking about how to get your code bases agent ready is creating agents.mmd that give it that are lightweight enough and give the agents a good kind of map of the codebase where to look at mapping out that is every agent now at your uh at your fingertips has a slash in it that gives you a good draft and then you keep itering those are living documents that should evolve as the agent makes mistakes and works.

05:29 And that's where I see a lot of teams that have invested in developer experience actually gain a lot because once they have good documentation and onboarding for a codebase that is read by the agent as well and that can help the agent as much. Next up you have skills. VS Code has always messed heavily in accessibility. So one skill we added early on was infusing all the accessibility best practices and how we think about accessibility into a skill that everybody will use and that previously was one person that had to

05:59 then got get pulled in or give feedback and now that skill was reviewed and is maintained by the kind of area owner of accessibility. So to bring that that accessibility that short work everywhere into everybody's hands. So think how you can encode your expert knowledge to unblock this like few users that might be always asked for this input into a skill.

06:22 And lastly once you do it right for us in VS Code the litmus test is can I as a PM effectively VIP code in the VS code repo which I do um and how is that accepted by engineering? How much work does it cost? Is it out of the box working? Is it for me working locally? Or is it eventually breaking when I ship it? So, a lot of these quality improvements will show in the end that you can allow everybody to be more effective in the repo.

07:01 Another improvement, and that's a massive one, is we switched to TypeScript Go. And if you've been in any developer codebase that takes some time to build or run linting or run CI/CD, you've noticed it's already annoying. But developers find other ways. They can review a PR. They can context switch into something else while runs and then do something else.

07:23 But once you hit that with agents, you have 10 agents, 20 agents running and hitting the same bottleneck, these any slow part of your CI/CD will compound. And for us, Typescript Go was a 10 times improvements in our builds, which at the scale of automatic PRs running with agents and them fixing getting needing that CI/CD loop for feedback to improve the code is a massive improvement.

07:55 And that leads to me uh filing big PRs. That was beginning of the year to get the ask questions tool in which I believe we really needed and then bringing it to the team to show to let them try it out in the product. Within a few weeks we had a perfectly polished and designed experience based on my initially very scrappy PR. But it unblocked that this these discussions the participation of everybody else because we had the ground set laid and the idea and the experience was clear.

08:30 Next up, when agents actually work on UI, and we all had this, that the agent gives you a UI and it says it looks perfect, then you open it and it's all misaligned. Even with SVG, even with the latest models, you still get it. And it's like the number one thing I keep telling everybody now these days. If your agent cannot use your application, your product directly to get this feedback loop of is it all working, then that's a really big investment that pays off every time you work on UI.

09:02 And there's two areas we invested in. One is the component browser, which is an automated build that runs every time VS Code changes something. Let me open this in this issue as an example. In this case, we added a back button to our customization screen. And because it's changing wide, we actually have this automated process in the back that takes screenshots of every component in VS Code and points out the differences.

09:33 It allows us to detect unexpected changes like we're changing one component as the ripple effect of something else moving or some other icon disappearing. and allows you to quickly review a PR as well because you also get directly a view and what the PR is actually doing and that's what we ask every developer before we ask them to at least attach a video or screenshot on what they have been implementing so we don't have to run everything from scratch and we can quickly assess and and give feedback and this is now

10:02 automated with the component explorer next up we have the self-correcting loop because in VS Code we run uh a web browser and open a website. Basically, that's what VS Code is. It's all HTML. We can actually use Playright to automate a lot in the browser. So, we have a /launch skill that would just open up VS Code and give it a specific scenario to click through in VS Code.

10:29 So, it's a really powerful way to diagnose issues because you can also get the locks along the way, but also validate any fixes you did before and after without you having to click through it. So you can just hit launch move on to the next task and then come back once the agent verified it's worked because it can access your application and to call out this is nice because we have playright and we are a web app uh running in electron but there's many other solutions out there one of my favorites is Xcode MCP because on

10:57 my site I work on iOS apps and once you see Xcode MCP just creating in VS code without Xcode involvement just making an iOS app opening it up clicking through it, taking screenshots. These moments are magical and give you a really tight feedback loop as well. So there's other tools available for Android as well to make this work. Let's talk about code review, which I think is the next big investment.

11:22 We know in the past when we enabled code we on GitHub the automated GitHub copiloted one we're not convinced yet but over the next months and weeks they have massively improved and now it's a mandatory uh review every time somebody opens a PR and there's more improvements coming actually dial in the uh the effort that a code review does depending on how much risk you have in this repo so you have a low medium high effort for code review so you can dial in kind of the cost benefit as well but for us is if a review after

11:54 review is done, humans will not even look at PR until all the comments are resolved and and addressed. Okay, now we're holding we're shipping faster. We got to hold quality. And a big signal for us because we're on GitHub is of course issues. Uh who has filed a GitHub issue? Who has had problems in VS Code before? Yeah, cool. Kind of similar. Awesome.

12:21 This the last crowd had less hands up. So it's good. Um, yeah. So it's a massive, it's a really good signal. And as I said, we get sometimes even better issues now because of AI and more people with different language backgrounds being able to file very good issues. But we also get a lot of issues. And previously engineers actually triage them manually.

12:39 And now we have AI doing the filtering to get out spam enriching and translating in the end assigning area owners for all the issues to raise the signal to noise quality of this. We also add a human layer to that once you are in issue because the agent will make mistakes and we need that feedback loop when it makes mistakes. We create a Chrome extension that allows you to fix duplicated issues as well.

13:06 So we have the agent doing the work and we have a feedback loop where humans can still fix it that we can bring back into the into the agent work up front. That's a really important thing to think about as you have agents doing the work that you get the human feedback as well. The other input we have is a lot of error stacks. Every time there's an exception and most apps will have this you get an error stack reported and we actually don't use any external tool for that.

13:34 It's all in our own data and we built all the tooling around it and it's mostly we already had it. We already had an ML doing sorting on top and this is now actually done by a mix of criteria. So one is we collect the raw telemetry. We filter it down and that's like 51 billion per day. We filter it down to just error stacks that actually have the full stack.

13:56 Then we group it and bucket it. That's more of the fingerprinting exercise. And then in the end after we all of that we have 10 issues filed assigned to specific area owners that a PR is autocreated to actually already try to fix the problem. This is our error page where we see all the stacks including how many hits and how many users affected so we can quickly react to anything that's coming.

14:23 And once we have this working we get a issue filed and a PR automatically opened. This is the issue. is the error with some initial diagnosis and then the PR got opened with and the agent actually goes in and finds out why does this happen who was responsible in this case um added a change that was connected to this to find f find symbols and then it figured that the on cancellation request wasn't in the RPC protocol so the agent already opened the PR we could just merge it and that's a really great automated way you

14:58 don't care about errors you just want to get a more stable code base. So all of this automated the triage initially mix of agents and deterministic systems then the issue is filed handed off to an agent in the PR multi- aent system and in the end we have fixing fixes landing mostly automatically with some human oversight to still approve the PR. There you go.

15:25 So previously VS Code did yolo releases in a way that it was release day everything was stable. We did our testing and we shipped the release to 100%. We just opened the floodgates but because of agents we want to derisk that process. We should have probably done it earlier. So now actually we do stage rollouts and as we do the roll out we monitor error locks and issues and everything else.

15:47 So just like good citizens in a in an web application, we now apply the same because rollbacks on an install app are so expensive. Okay, now we have this shipping faster, shipping at quality. How do we actually do this in a in a speed that we actually learn faster, which is the which is a bigger challenge with AI? You actually want to build a better product faster.

16:12 So one of the examples is our VSC bench. We're shipping an agentic product. So it's important that we have our own evals of our product always at hand and being really easy to extend. It's all in the GitHub issue. It's all in GitHub. When we file an issue to add a new scenario, it's actually an agent picking up from a template and bootstrapping the scenario for you.

16:32 So we make it as easy as possible to bring in all the developer scenarios that are coming up in issues through customer conversations through all these channels to bring them back into our emails. So we then start hill climbing actually against them and then com also confirm that it's all improving offline and then an online experimentation. So classic product life cycle for um for us and we see these interesting things in our stack.

17:00 Understanding your evals is critical. One experiment we did that you can read about on our blog is that we gave a very simple eval scenario of just write a file with hello world just testing out our eval harness end to end and how different models react. The same five character file took the most expensive model 70x more tokens and that was not the high reasoning model.

17:26 You can read more about unblock. We actually don't talk about the models, but just that insight in how your models react with your harness is extremely is a very important exercise. Next up, we have prototyping, which I do a lot of because I noticed that I don't actually want to land most of my work into VS Code PRs itself. It's a great way to start a conversation, but most often I want to have a quick conversation because now we're meeting on a daily basis where we just want to have very quick feedback loops on here's

17:58 the idea, here's how it could look like, what if we do this. Next day you come back with updated uh prototypes and you just keep having this conversation. So much quicker feedback from a month release cycle to weekly to daily sprints where you just keep working on problems and prototypes unlock those deeper discussions and how the experience should look like.

18:21 It also means that we have the durable ownership in internal work as I mentioned daily sprints um smaller squads, smaller work streams that really prioritize and drive smaller sets um of areas to move faster. So try these things, find out where your bottlenecks are to actually ship higher quality, ship faster and learn faster and find fix that next bottleneck.

18:46 And don't just tune how you work with agents, but tune where how how do they get feedback? How do you build these loops? And then figure out after you have these these success moments, how can you improve that feedback loop for the next moment of what's blocking you to move even faster and it's a constant iterative cycle across all these cycle ones. That's it. Come by code booth. Happy to chat now. Otherwise, happy coding. >> [applause] [music]