← All transcripts

Move Fast and Don't Break Things: Scaling Databases for the AI Era — PlanetScale Transcript, AI Summary & Key Points

AI Engineer · 2 days ago · Science & Technology · 19:07 · EN

Watch on YouTube

AI Summary

Databases can scale for the AI era without sacrificing uptime — isolation, redundancy, sharding, back pressure and decoupling handle both reliability and the traffic surges that AI agents cause. A JSON VSchema lets an agent design a sharding plan for Vitess, Traffic Control sets resource budgets so a surge of agent traffic degrades gracefully instead of crashing, and database branching, deploy requests and one-click reverts let agents change schemas safely. Good developer experience turns out to be good primitives for AI.

Key Points

  • Facebook's 'move fast and break things' phrase is twenty years old and predates AI agents writing and deploying code — shipping fast matters more than ever, but it can be done without breaking databases or infrastructure
  • AI demand has doubled, tripled or 100x'd daily active users for many companies in the past six months to two years, leaving infrastructure struggling to keep up
  • Isolation: keep the data plane (the critical database) separate from control planes, observability pipelines, analytics, apps and MCPs so that deploying a bad change to the control plane cannot take down access to the database
  • Redundancy: a primary node with replica or follower servers in different data centers or availability zones, because at thousands of servers a once- or twice-a-year failure becomes a weekly or monthly failure; application servers are easier since stateless workloads autoscale via Cloudflare Workers, Vercel functions and AWS Lambda
  • Sharding: spread data and queries across many servers with a proxy in front (VTGate for Vitess), used for 10 terabytes to a petabyte where a single node fails; PlanetScale does this through Vitess for MySQL and Neki for Postgres, and each shard still needs its own primary plus replicas for self-healing
  • Back pressure: MySQL, Postgres and SQLite crash when overloaded — too many queries, too much CPU or exhausted RAM — so design the system to deny requests and fail connections above a resource threshold, keeping part of the traffic healthy while gracefully degrading the rest
  • Decoupling: combining Postgres OLTP, analytics, queuing and caching in one service adds complexity, breaks failure isolation (one service going down takes everything down) and creates resource contention, so scale compute for each part independently
  • GitHub's traffic grew roughly 2-4x on top of an already gigantic load, attributed almost entirely to AI agents, especially over the holidays

Tools & resources

6 items

ANo. 4972
AIAINotes.us Tool

AWS Lambda

aws.amazon.com

AWS Lambda is a cloud backend function used in the described system to connect an Alexa skill to an application.

Mentioned in
2 videos
Kind
Other
DNo. 5246
AIAINotes.us Tool

Database Traffic Control

planetscale.com/database-traffic-control

Database Traffic Control is a PlanetScale feature for enforcing real-time resource budgets on Postgres query traffic. It matches workloads by query pattern, application name, Postgres user, or custom tags attached through SQL comments, then limits server share, burst capacity, per-query usage, or concurrency for each traffic slice. Budgets can run in warn mode to observe violations or enforce mode to block queries that exceed limits, with separate warning thresholds available for enforced budgets. It is designed for incident response, priority-based traffic shaping, protection against AI-agent traffic spikes, tenant and regional isolation, and safe deployment rollbacks. The feature is available for PlanetScale Postgres databases through the dashboard, API, and CLI, and integrates with PlanetScale Insights for finding problematic queries.

Mentioned in
1 video
Kind
Other
NNo. 5245
AIAINotes.us Tool

Neki

planetscale.com/blog/announcing-neki

Neki is PlanetScale's sharded PostgreSQL product, built by the team behind Vitess. It uses explicit sharding to distribute PostgreSQL data and queries across multiple servers for demanding workloads. Neki is not a Vitess fork; PlanetScale says it was designed from first principles for PostgreSQL and was developed with large-scale design partners. The product was announced as available in Platform Preview.

Mentioned in
1 video
Kind
Other
PNo. 5244
AIAINotes.us Tool

PlanetScale

planetscale.com

PlanetScale is a managed cloud database platform for Vitess/MySQL, Neki sharded Postgres, and PostgreSQL. Vitess and Neki support horizontal scaling through sharding, while the platform provides connection pooling, read replicas, online schema changes, branch-per-environment workflows, automated backups, failover, and online cluster resizing and resharding. Its database workflow includes schema branches, deploy requests, revertable changes, and Database Traffic Control, which enforces query-traffic budgets to limit runaway queries and unexpected load spikes. PlanetScale exposes these capabilities through its platform tooling, APIs, CLI, and an MCP server for controlled agent-operated database changes. The service offers hosted deployments and a bring-your-own-cloud option through PlanetScale Managed, with deployments on AWS and GCP. The homepage states that multi-region deployments have a 99.999% SLA commitment and single-region deployments have a 99.99% commitment.

Mentioned in
1 video
Kind
Other
VNo. 0598
AIAINotes.us Tool

Vercel

Open source · vercel

Vercel is an application and agentic infrastructure platform for deploying and hosting web applications, marketing sites, platform products, and AI agents. Its infrastructure includes global delivery, deployment environments, serverless functions, fluid compute, a web application firewall, durable orchestration, sandboxed environments, and an AI model gateway. For hosted platforms, Vercel provides tenant isolation, domain management, custom SSL certificates, and preview URLs.

Mentioned in
9 videos
Kind
Other
VNo. 5243
AIAINotes.us Tool

Vitess

vitess.io

Vitess is an open-source, cloud-native database system that provides a MySQL-compatible interface while extending MySQL for larger-scale deployments. Its built-in sharding routes application queries across shards without requiring sharding logic in the application; the video identifies VTGate as the intelligent proxy that parses MySQL queries and sends them to the appropriate shard, with sharding plans specified in a JSON VSchema file. Vitess also provides query rewriting, caching, connection pooling, materialized views, cross-shard messaging, live resharding, background schema changes, and automatic detection and repair of primary failures. It is a graduated Cloud Native Computing Foundation project.

Mentioned in
1 video
Kind
Other

Links mentioned

🔒 Full analysis locked

Unlock more videos and the full analysis

A credit unlocks one video's full analysis for good — the build steps, the tools and how each was used, the methods behind every use case. Pro opens the whole library instead, and raises how many videos you can analyse a day.

Unlock full analysis — free

Transcript

Searchable transcript of Move Fast and Don't Break Things: Scaling Databases for the AI Era — PlanetScale — AI Engineer (19:07). Search for a phrase, then click its timestamp to jump straight to that moment in the video.

Captions sourced from the original video on YouTube, published by AI Engineer. The video, its captions and all related intellectual property remain the property of their respective owners; AINotes claims no ownership. Provided for research, accessibility and search — see the Transcript Notice and Copyright Policy.

00:12 Hey everybody, how's it going? Hope you're having a good last day of AI Engineer. Uh what I'm going to talk about today is how to move fast and not break things. So the the phrase move fast and break things. You've probably heard this before and what this uh refers to is it's a phrase from early Facebook engineering culture days, right? And the idea was we want to be able to move fast and ship fast and not be constrained by, you know, databases going down or breaking certain features because the priority is there's so

00:40 much value in shipping fast. And this was 20 years ago, right, before we had AI agents writing code for us and shipping things and deploying things on our behalf. So, if anything, the idea of doing this is more true than ever. But in my opinion, we can actually achieve both, right? I want to be able to ship fast, build great features, get things out to my customers, but also not break the databases, the infrastructure, the application servers that are powering what I'm doing.

01:06 And we've probably all seen something that looks approximately like this. And I'm not calling out any particular company. There's a number of companies right that due to the kind of insane scale that we've seen over the past even 6 months year two years right where we've had like daily active users double triple 10x 100x a lot of that powered by like the demand of AI agents um where the infrastructure struggles to keep up and some of this is simply because it's a very hard problem right this is not meant to call

01:34 anybody in particular out it's hard to scale things to tens of millions or hundreds of millions of users but what we'd like to have is great products that scale well for lots of users and have our status pages look a little bit more like this one, right? Three or four nines of availability uh maybe five of availability where our databases are healthy, our applications are healthy and ultimately our users are happy, right?

01:59 So this is kind of the idea behind what I want to talk about which is a little different than probably most of the talks you've been at. A lot of the talks are probably talking about inference or specifically how to do things with agents or different kinds of automations. This is more of an infrastructure talk, but how you can do that and ship fast both with the demand of AI agents as well as the you being able to use AI agents to make all of these processes more optimized and smoother.

02:25 So, there's three parts to this that I want to talk about. The first one is how do you even build systems regardless of the number of users you have that aren't going down all the time? and then tying that into well okay if you're going to build the next hit application that's going to scale to millions of people using it how do we scale that using these same principles and then finally how can we take these two things and let AI agents do them for us or at least in part work with them in doing this so I should

02:54 probably introduce myself a little bit too why I'm talking about this is I'm Ben and I'm from planet scale and we do databases this is our whole thing and we work with a lot of these AI companies that are scaling extremely fast. We work with companies like cursor for example and power databases for them. And so this is something that we as a company take very seriously.

03:14 How do we scale and support companies that have millions of users uh on their platforms and some of what I'm talking about comes from a article that we published last year about kind of our philosophy of how we operate databases. I would encourage you to go Google that and read that. Um but a couple of these principles let's go over some of these. So the the first and one of the most important is isolation.

03:36 Right? When you are running infrastructure, one of the things you want to make sure you do is have a setup where if something fails, it doesn't take down the whole system, but every component is very isolated. So what I'm showing here is data plane. That basically means where are your databases? Where does your most critical data live that you need to be using for your users to log in and interact with your application?

03:57 But there's lots of other components, right? There's your control planes, there's your observability pipelines, there might be your analytics frameworks, there's your apps, there's your MCPs, and there's all these different components that can potentially fail. And in fact, assuming you're using a good database platform, the database is hopefully low likely to fail, but these other things you're shipping and deploying to more frequently, so are more likely to fail.

04:19 So we want to make sure that we've designed things where if I deploy a bad ship to the control plane and take it down, nothing else is impacted and can still access the database, right? And anything in my data plane. So this is one of these kind of core principles that we take very seriously. And another thing is redundancy. And if any of you have worked with databases before in the past, you've probably seen a diagram that looks something like this.

04:44 And similar things apply to your app infrastructure and your other infrastructure. Although in the case of application servers, usually redundancy is a little bit easier to solve because it's a stateless workload, right? There's lots of sort of autoscaling cloudflare workers and versell functions and AWS Lambda that can scale and have redundancy very easily because it's not storing state.

05:05 But in the database world, you have to store state reliably. You can't lose a single bite of data. And so how do you do this? Um, a very common approach which costs more money and is more complex but it's the way that you scale reliably is by having a primary node. So this is where most of your traffic goes to. But you always have replica servers or follower servers standing by in different data centers or different availability zones ready to take over because the fact is servers do fail sometimes.

05:31 Uh, and on a small scale when you've got an app that has a hundred users and you have one database server, you might only experience a failure once every couple of years. But at a large scale when you have thousands of servers a once or twice a year failure becomes a weekly or monthly failure. And so you need to have systems in place to deal with this very well.

05:53 Okay. So we've talked about uh what the principles of of scaling reliably are or of um having reliable systems are. But then the question is what I've just told you here we could write a whole book on. Right? So there's lots more that you know there's people who spend their entire careers focused on this. So, we're going a little bit quickly through it, but these principles apply no matter what point of this user growth chart you're at, right?

06:18 You could be at 100 users and in this case, uh, it's still good to have redundancies, to have isolation, to have failure even for that smaller set of users. And you want to apply those same principles when you are all the way up here. I don't know if you can see that line very well, but at millions of daily active users. But there is a point where you actually have to incorporate different technologies to scale efficiently and effectively.

06:41 So, uh, let's talk about how we scale those systems for a few minutes. And probably the part you're most excited about is the AI agents part, which we'll get to at the end. Um, so one of the key things, and this is what we do a lot of at Planet Scale, is sharding. Um, I just realized my mouse is on the screen, so let me get that off the screen here.

07:00 See if that will go away eventually. Um, so database sharding. What is this? What you do is instead of relying on that single node, which works fine when you have 10 gigabytes of data, it doesn't work so well when you have 10 terabytes of data or 100 terabytes of data or a pabyte of data. So what do you have to start doing? Well, what you need to do is spread your data and your queries out across many servers.

07:23 And so that's what all of these shards down here below are. And typically you put some kind of proxy and we'll talk more about what VTGate means later in between that intelligently handles distributing all of the user requests or the query requests across all of these servers. Uh and this is exactly what we do a lot of for some large customers at planet scale.

07:44 We do sharding through vitess for my SQL and niki for postgress. Um and so this is a very fundamental way of doing sharding. But the other thing is you don't just use sharding by itself. If we look at this and we see for example shard A, within that box, you actually would want a redundant system. So you'd have a primary and several replicas in there.

08:04 So if one of those servers fails, that single shard can do self-healing and get back online very quickly if there's a problem. So this is one of the ways that we see and is probably the most common and performant way to scale databases. But really the same principle of hey let's take queries and data and spread them out across many servers, many nodes, many drives uh applies to lots of other areas of the stack that you're building.

08:29 Back pressure. So this is another one that is sort of a mix of being a part of scalability but also a part of reliability. And what we mean here is there's a lot of systems like MySQL and Postgress and SQLite, these very popular databases that people run that don't do a good job at handling being overloaded. The moment that they have too many queries or using too much CPU or they run out of RAM for the cache in many cases, if not taken a lot of care for, they'll just crash and now your server is down and now all of

08:58 your users are impacted and they can't use your service that the database is powering. What back pressure is is it's designing a system to be able to push back and say when I detect that I'm above a certain threshold of resources I can actually say I want to push back and start failing requests denying query uh denying connections right so that I can keep the existing users or some of the traffic healthy while gracefully degrading some of the other traffic and then the rest of your stack has to be built to support that.

09:25 So, this is a principle that applies in lots of areas of the stack, but in the database, it's arguably the most important. Uh, let's see here. The clicker. There we go. And then the last component of this that I want to talk about before we move into how to do this all with AI is decoupling. Uh, there's a lot of people that you'll talk to and when it comes to databases will argue that you should probably have all of your services combined into one because there's some convenience to that.

09:52 your Postgress database, your OOLTP workloads combined with what you're doing for analytics combined with your your queuing and your job system, your caching. There is convenience and some technologies try to do this where we put this all in one. But that does one add a lot of complexity, but it also breaks some of the fundamental principles from earlier like failure isolation, right?

10:12 If one service goes down, that actually means everything is going down. And then you also have to deal with contention of resources. So if you suddenly have a burst of Q jobs and you need to scale up your cues, but maybe nothing else really needs to scale, it would be nice if we can independently scale the amount of compute resources that we've given to each of these parts of our backend service powering your favorite application.

10:35 Right? So this is not a complete list. What did we go through so far? We went through sharding. We went through back pressure. And now we're talking about decoupling. There's many other aspects to scalability. And again, you could spend your entire career becoming an expert in scalability. Uh, but this is a lot of what we think about a lot at planet scale because our customers rely on us to be reliable and to scale.

11:01 So, finally, let's go to the last part of this. We let's say we have all of these principles in place. We are doing things reliably where we have fault tolerance. We're doing things that are very scalable. But these days in a lot of organizations and we talk to a lot of people that feel this both at the small scale and at very large organizations. How can we allow agents to whether in part or in full have control over what is arguably the most important infrastructure of your company, right?

11:31 The databases, the cues, the things that are powering every other layer of your stack. Um, and this is something that can be a little bit intimidating because it's pretty much the standard these days, right, to write code with AI. In some cases, people are not even writing any lines of code anymore, right? AI is actually writing all of your code, reviewing all of your code.

11:48 Um, and you're just sort of guiding these agents to doing that. But it is quite different to say, I'm going to now also give my agents the ability to deploy changes to the database or make changes to how things are configured there or to shard my database, right? because that's something if it gets wrong, it's generally not as simple as clicking a button to roll it back, right?

12:08 It's taking down your entire application and every user notices. So, we want to do this effectively and safely. So, I'm going to revisit a couple of the concepts that I previously talked about and think about how do we use agents to solve this. So, going back to sharding, what I'm going to talk a little more specifically about is Vitess and this is an open- source project.

12:27 uh we operate the test databases for our customers and it works with my SQL to do sharding. So what you can see here is like all of those shards down there those would be my SQL databases and the VT gate that's the proxy that is basically an intelligent proxy that takes my SQL queries parses them figures out which shard they need to go to and routes all of those requests in a very sophisticated way.

12:49 So the software is fairly complicated, but the interface that we present for specifying how do you want all of your queries and your data sharded is actually quite simple. It's what we call a v schema and it's literally a JSON file where you say here's the column I want to shard on and here's the way I want to do sharding and you can specify certain tables to go on certain shards and have lots of control over where and how things are configured.

13:12 And the great thing is, as we all know, AI agents are excellent at working with things like text files and JSON files. So you can give your agent and say, "Hey, I'm struggling with scalability of my database. Here's my schema. I want you to come up and help me prototype a good v schema." Right? And once you've developed and designed something that it has thought through and made decisions about the most efficient ways to shard when you are on a platform like planet scale you literally will plug this in and you make a

13:39 deploy and you you can we give you an interface for creating new sharded systems and scaling up your database and a lot of this can happen with either the help of AI or letting AI take some control over these processes uh given our API and our command line interface that we run. So this is one of those things right very complex infrastructure in reality under the hood that is scaling your systems but you provide the right interfaces for it like simple configuration files and those things become much more accessible to

14:05 any company trying to scale. Um, another thing, so this is getting into the uh concept of back pressure that we talked about earlier, right? So when a database is overloaded, and this, as you might imagine, probably happens all the time at some of these companies, right? I think there was a a tweet a couple of months ago from GitHub that basically said like their traffic had like two or three or 4x from what already was an extremely high load.

14:30 And pretty much all of that was attributed to AI agents, right? So they already had gigantic infrastructure because they're the most popular place in the world to develop. and then they had multiple uh uh traffic increases especially over like the holidays right due to these uh overloaded systems. So how do we deal with this? One of our solutions and the way that we like to think about it is we give users a system called traffic control which allows you to essentially categorize and tag all of the traffic coming into

14:58 your database and then you create resource budgets to say hey this segment of my traffic cannot exceed this amount resources of CPU usage or of backend processes and again we do graceful degradation when that happens. So what I'm showing here is this is a graph of I have a budget set up um to say don't exceed these amount of resources and warn if I do start exceeding these resources.

15:19 So it's showing me uh this is the amount of queries that are exceeding your budget over time and we're we're we're sending warnings back to the client like hey I'm overloaded you need to slow down and we can even make it more strict where we actually kill those queries that are executing. Right? And this is one of those things that like, okay, it's not ideal to have to deny requests, but it's much better than your whole application going down and being offline for all of your users.

15:45 Uh, and then finally, one other thing I I want to talk about, and I know I'm showing this is the planet scale UI here, but this is not even so much specifically about planet scale, but it's about the philosophy of how we do these things with agents, which is most developers are used to using a gitlike flow when you're building. And this applies to your agents as well, right?

16:04 You have a main, you create a branch, you make code changes, you do merge that in, and you deploy that dynamically to wherever your application is deployed. What we have done, and this was long before agents even existed, is we wanted to take that same UX and mirror that for your database because a lot of the times when you're building new application features, you're also making changes to your schema.

16:25 You're adding tables, you're adding columns, you're deleting things, you're you're adding indexes, you're doing all of this stuff to your database to keep it in sync with your codebase. So what we give you powered by that same vitess earlier is the ability to branch your database schema, make changes to that schema in an isolated environment and then merge that back into production with what we call a deploy request.

16:45 All with no downtime. And not only that, but actually the ability, you might see here at the bottom, the ability to, if you deploy a schema change and you see a problem, we give you the ability to oneclick revert that go back to the old state of your database without data loss. Uh, and so this is something that was great for humans, but we also have the APIs and the CLI to be able to let agents automate this.

17:07 So when they are doing things and automating code and branching and making pull requests and doing code review, there can actually be a mirror database state that is working in sync with all of those things, right? And that can either be fully automated, it can be semi-automated. Obviously, many of us are working in a human in the loop fashion where humans are still reviewing all of these changes that AI agents are making, which is a good thing.

17:29 Um, but we have the primitives in place for building these uh systems. And there's a number of other things that I didn't get to to talk about this, but kind of the the ending point that I want to make is we're a company and I know a lot of companies out there like to focus on the developer experience. And when I say developer experience, I don't just mean like light and dark mode switches and good keyboard shortcuts and all of this, but I mean the actual experience of, hey, if you're a developer that's responsible for

17:58 your company's database or responsible for some piece of infrastructure, do we give you the tools that you need to not mess up to do deploys safely to make your database faster to get better performance? Uh, and we've been building these things for years. And it does kind of turn out that when you really focus on that core of the developer experience, what you actually end up with is also very good primitives for AI being able to work with a database safely and reliably and helping it scale.

18:23 Uh, and so because we expose a lot of these things through our MCP server, through command line interfaces, obviously there's plenty of debate on which of those is better when working with your AI agents, but because we expose those things to both parties, it's actually a great experience for operating your database and working well with AI agents. So that is all that I have to say. Thank you guys for coming. I'm Ben and have a great rest of the conference. >> [music]