Avoid cascading failure by reducing hazards, strengthening components, and reducing interconnection, while consciously balancing resilience benefits against added cost and complexity.
AKS is a Kubernetes platform used as an example of an operational model that differs from Amazon EKS.
Amazon EKS is an AWS-managed Kubernetes environment. The video uses it as an example of a Kubernetes platform whose operation differs from other Kubernetes environments.
Amazon Web Services (AWS) is a cloud computing platform and subsidiary of Amazon that provides on-demand computing resources and services including compute, storage, databases, networking, analytics, machine learning, and developer tools. AWS offers these services to individuals, enterprises, and governments on a pay-as-you-go basis.
Google's cloud computing platform, providing AI and cloud computing services for security, data management, and hybrid and multicloud environments. In the cited video, GCP is discussed as a source of rented GPU spot instances when a company’s owned or contracted capacity is insufficient.
Kubernetes (K8s) is an open-source container orchestration system hosted by the Cloud Native Computing Foundation. It manages containerized applications across multiple hosts, providing mechanisms for deploying, maintaining, scaling, and scheduling workloads. Applications can be configured with YAML and deployed across hybrid-cloud environments; the platform is also used for production operations involving storage, databases, networking, and security hardening.
Microsoft Azure is an open cloud computing platform and enterprise infrastructure service used to deploy and operate software. It provides infrastructure as a service, virtual machines, databases, containers, serverless computing, Kubernetes, storage, networking, monitoring, security, analytics, and developer tools. Azure also hosts AI-oriented services for models, machine learning, applications, and agents, and connects AI, data, business context, applications, and agents across an enterprise.
Searchable transcript of Understanding Progressive Collapse: How To Avoid A Cascading Failure — InfoQ (48:06). Search for a phrase, then click its timestamp to jump straight to that moment in the video.
Captions sourced from the original video on YouTube, published by InfoQ. The video, its captions and all related intellectual property remain the property of their respective owners; AINotes claims no ownership. Provided for research, accessibility and search — see the Transcript Notice and Copyright Policy.
00:01 [music] [music] >> Um I'm talking to you today about this concept called progressive collapse, which I came across while doing research for my latest book, Blame Pluck, will be out later this year. Details to follow at the end of the talk. Um uh but I came across this while looking at the wider concept and wider space of resilience engineering and lessons and things that we can learn from top picks and from domains that are not around computing.
00:33 And I learned about the story of Ronan Point. Uh this is a tower block um that was shortly after it was opened in 1968 suffered a partial collapse. Uh it was not a good time. This was in Canning Town, actually not that far from where we stand today. Um you can get a sense of the scale of what happened uh in this wide shot. Only four people died. I say only four people died.
00:57 It was a miracle it wasn't more. The only reason more people didn't die was because of the hour at which this particular collapse happened. It was quarter to six in the morning and most of the affected rooms there's basically corner gone. You can see the mirroring sister block on the right-hand side here. And most of those rooms that collapsed were living rooms.
01:16 That early in the morning very few people were up. We're going to be talking about what happened to Ronan Point and this concept called progressive collapse, which kind of comes from space of civil engineering. And also going to taking a look at how we can take that metaphor and use that to reason about our own digital systems. When things go wrong in the distributed digital systems that we're all creating and we're all creating and hopefully trying to make more resilient.
01:40 And taking those lessons that we can learn from the building industry, what are the techniques that we can use to stop progressive collapses happening in our own system. So I'm going to take you into a bit more detail about Ronan Point, but also a couple of other examples of system failure, one of which I was intimately involved with, and one of which may have impacted many of you last year.
02:01 So, what happened Ronan Point? Uh these buildings were built quick. This is post-war. Uh although this is 1960s, there was still an awful lot of London, especially in the Canning Town area, that was still rubble after the Blitz. And there was a boom in population, we needed more housing, and so a lot of these blocks went up quick. We'll come back to the building methods used a bit later on.
02:24 Um it was early in the morning. Uh resident uh Miss Ivy Hodge got up. First thing any self-respecting British person does, especially in the late '60s, was make tea. She put like a little Yeah, she put a stove on. Went to heat up some water, and there was a gas explosion. The gas explosion was caused by uh just a faulty knob around the uh where the oven was connected to the main gas supply.
02:50 This was a survivable gas explosion, and I mean that quite literally in terms of It wasn't a great time for for Miss Hodge. She was blown across the room, knocked unconscious. This is what's left of her kitchen. She came to in a puddle of water from the saucepan that was blown off the top. This was not a big gas explosion in the grand scheme of things.
03:09 Still, uh not an ideal scenario for first thing in the morning. Uh you can actually see some of the problems start to occur beyond that Miss Ivy Hodge was in there somewhere. When you look, you can see through the side of the building now. This is where our problem gets a lot worse. This gives you a plan idea of what happened. So, you can see uh where the sink was.
03:26 You can see the gas cooker was knocked over. And then you see where the building used to be on the left-hand side, that her built bedroom was sheared off. What had happened was that the explosion blew out the outer wall, which happened to be load-bearing. The four floors above collapsed down, and that creates the concertina effect taking down that corner of the building.
03:49 So, you can sort of see where the explosion happened. You We we might assume that the uh the impact would have been even greater if the explosion happened further down the building. So, this is what's known as a progressive collapse, and a progressive collapse describes a situation where a small failure results in a significant collapse in the wider system.
04:08 The initial triggering thing in isolation looks small, somewhat innocent, but it ends up cascading to cause a much bigger impact. We're all familiar with this idea. I'm sure some of you have either done this experiment or seen versions of this where you can start off with a little small domino hitting a bigger domino, and you can, you know, I've seen examples where they're putting over knocking over dominoes that are like over a ton in weight.
04:32 The differences with this kind of system is that we can see all the moving parts, and we can reason about it. The progressive collapses that happen in our digital distributed systems are not as easy to necessarily understand. We're going to take a look at two other examples of progressive collapses that occurred. Firstly, we'll take a look at the US East 1 outage in October of last year, and secondly, a website selling used cars.
04:55 This is my fault, but we'll come back to that. Let's talk about what happened with AWS. Actually, out of interest, put your hands up if you were if you or any of your systems were impacted by this outage. Yeah, a surprising number of like UK banks were impacted by this, which doesn't really make a lot of sense considering there is an AWS region here, but we'll maybe come back to that later.
05:19 Fundamentally, what had happened was a subsystem of DynamoDB was tasked with updating DNS routes within the AWS infrastructure. This was to allow for things like provisioning and sort of load distribution and the like. This had an error, which we'll get to the details of later, and this caused DNS routes in Route 53 to be deleted. This in turn had a cascading effect because other services then that relied on this DNS route started to fail.
05:48 Network load balancers started going down, compute started failing, queues stopped working, and EKS also failed. This is all compounded because a lot of these lower order services in AWS are used by other services in AWS. When those network load balancers start going, all hell breaks loose. And so, a lot of the US East 1 services started failing. This in turn took customer-facing products and companies offline.
06:16 Amazon's own Alexa stopped working, Ring, Slack, uh what Snapchat, Zoom, Shopify, all suffered either partial or total outages as a result of these particular failures. People's alarms didn't go off, so they were late to work. It seems like a good excuse to have though if you are late for work. Oh, I'm sorry, a cloud region went down. People's smart locks would not unlock or would not lock.
06:44 Um I saw at least one person on Reddit claiming they were locked in their house. Uh more humorously, people who have all the money in the world can buy a $2,000 smart mattress found their mattresses were stuck in certain positions, were stuck heating or cooling as the case may be. And the one I like the most was if you have 600 bucks, you can buy a smart litter tray.
07:06 Then these smart litter trays also stopped working. Don't worry, the cats were not trapped. Right? So, we had these interesting kind of rippling effects that occurred. So, let's actually go a bit deeper into what happened, and we're getting this information from AWS's own incident report. The this component of DynamoDB that updated the DNS routes, there's kind of two main parts of it.
07:26 Think of a DNS planner, which worked out which DNS entries need to be updated, so they create plans, which kind of makes sense. And then you had a number of DNS enactors, and their job was to enact the plan. Naming stuff, it turns out, is quite easy. So, what happened was a DNS enactor would pick up a plan and apply those changes. There are multiple copies of the DNS enactor for resiliency, which really ended up causing this problem, but we'll come back to that later on.
07:56 The issue was on the particular day or particular day the incident was one of these plans took a long time to implement, much longer than normal. It was caused by a host of underlying issues. This plan had not completed, a new plan got created, a different DNS enactor picked up that other plan. It completed its work very, very quickly. Unfortunately, then the old plan completed and wiped out all the routes.
08:22 Um there is more detail here in the incident report. I'm glossing over a few details, but this was a classic race condition resulting in all those DNS entries being wiped out. So, what happened with the used car site? What about this kind of failure? Uh well, this was a website, it still exists. I was involved in many years ago. Was for helping you find used cars, motorbikes, caravans, and the works.
08:43 And it had historically been a bunch of sort of websites based around those particular verticals. We're in the process of combining all of those verticals into a single application. This is actually a classic strangler fig application. In fact, this is the case study from Martin's original write-up. So, our system, which is code-named Sauron, was running on 10 servers.
09:05 Um and typically, even at peak load, we would expect have something like between 30 to 60 concurrent requests at any given point in time. Some of that traffic was being served directly by Sauron. Other calls coming in, though, were being passed on to the downstream existing legacy applications. On the day the incident occurred, these nodes started running a little bit hot.
09:25 They went from handling between 30 to 60 concurrent requests to handling over 800. This is a Java stack, native threading, looks interesting we're back to green threads again. So, that meant each of those concurrent requests became an operating system thread, and so the entire CPU was saturated as it was spending all of its time trying to sequence between those threads.
09:47 This took the whole system down, and it went down quick. So, what was the issue? Well, we had all these requests coming in. It wasn't a denial of service attack, we eliminated that pretty quickly. Turned out the issue was that one of the downstream sites started behaving in a sub-optimal way. It would It would let you establish a connection, and it would just hang.
10:09 And it would just hang forever. Well, and we were timing out way too leniently. We were taking 30 seconds before we timed out those calls. This basically resulted in the connection pool that we were using to make calls becoming completely exhausted. So, now we had a thread pool issue. Because that connection pool was exhausted, any other calls coming in that were for motorbikes or for jet skis were unable to get a workup, so those calls themselves were also blocking.
10:39 This was made worse because when the website hung, what would people do? They hit refresh. So, they would just hit refresh, and then the pages wouldn't load, and they hit refresh again, and the pages wouldn't load, and they kept hitting refresh. And we weren't timing out on the calls coming in. So, not ideal. There were a bunch of different things that we did to resolve this issue going forward.
11:02 I'm going to take you through some of them, and the ones that we missed at the time. So, these are kind of three examples we looked at. We've got Ronan Point, we've got AWS, and we've got the used car site. So, how do we mitigate for a progressive collapse? Right, we can't eliminate the possibility, but what things can we do to mitigate against these things happening in our own systems?
11:23 Now, I want to kind of preface this by saying that focusing on a single root cause is highly problematic. Note to whichever marketing team works with one of the vendors upstairs that keeps talking about finding the root cause. No. Bad vendor, right? Thinking about a single root cause is deeply problematic. There should be at the very least you should use the word root causes.
11:45 But in general I try and avoid this because it makes us think overly simplistically about the world. You might be in a region which is troubled by forest fires. I lived for many years in Australia. You might think there was a big fire, not good. It was caused because somebody dropped a match. Well, that's the root cause. Let's stop people dropping matches.
12:04 Now that might be a good idea, but it's not by itself sufficient. In Australia for example, you're not allowed open fires in bush areas. You're not allowed to use things like fireworks either because well duh, they cause fires. So doing that is useful, but it's not by itself sufficient. And so they do things like they clear the brush and the undergrowth.
12:25 They do controlled burning to reduce the kind of the the load within as well because they recognize that although they can try and remove the thing that maybe causes the fire initially, you can't remove all causes of fire. Lightning is a thing. So other things are needed to be done to reduce the impact of a forest fire if it occurs. And the problem with this is that a lot of this viewpoint about finding the root cause still permeates our industry.
12:52 I often think that IT is still stuck in the world of traditional safety management. And this is a fixation on the idea that we just review reduce adverse events and that makes things better. We try and stop anything bad from ever happening. The problem is as our systems get bigger and our systems become more complicated, we cannot stop all the bad things happening and in any case a lot of the bad things are entirely out of our control.
13:20 This is why in the world of resiliency we've gone from this world of traditional safety management and have moved over to the world of resiliency engineering which is a big topic and it goes far further than I've got time to explain to you today. But in a nutshell, we can think of resiliency engineering as accepting that things can go wrong and still try to mitigate them.
13:36 However, accepting that and trying to make sure as many things go right as possible. This is a very different mindset. So, accepting that things might go wrong, but knowing how to mitigate those things when they occur. So, in that spirit, what can we do to mitigate a progressive collapse? Now, of course, at this point, I go to your friend and mine, NIST.
13:59 So, I popped over to NIST, uh which is replete with very interesting PDFs and all kinds of topics. And we have this report on practices reducing the potential for progressive collapse in buildings, not in distributed systems. That's the only flaw with this particular report. Um it is 216 pages. I did not read all of them. I put it into Notebook LM. I asked Notebook LM questions.
14:22 Not going to lie, but you can download it and there's actually lots of really interesting advice in there. But I was basically able to synthesize kind of the three key ways that you can mitigate for a progressive collapse. And I'll talk to you about how those were used in the context of buildings. And I'll also talk to you about how those three types of mitigations can be applied in our digital systems.
14:45 So, we're going to look at the importance of reducing hazards, of strengthening components in our system, and reducing the interconnection in our systems. So, then we're going to look at each of these three things in turn. And so, if you're going to try and mitigate progressive collapse in your own environment, you probably want to do something in all of these areas, if you can.
15:06 So, let's talk about hazard reduction. So, in the wake of Ronan Point, what could we do to reduce hazards? Well, immediately in the aftermath of the incident, it was suggested to stop allowing gas to be installed in high-rise buildings. Uh a lot of these buildings were going up, and there's over a thousand buildings that have the same building system as this across the country.
15:28 So, just don't install gas. That would remove at least that inciting element. In those situations where gas was already installed, it was suggested that we should improve ventilation. This avoids gas building up, but also makes it more likely that other people are going to smell the gas and alert the authorities. Any of you who have spent any time living in the UK know ventilation and where we live tend not to line up very very well, but you know, we all have damp issues around the winter, don't we?
15:55 But nonetheless, it seems like a sensible thing to do to reduce the hazard of a gas explosion. What about AWS? Well, we can look into at least what AWS talked about doing again from that incident report. And very sensibly, they said, "We've clearly got an issue with our automated DNS management. So, while we work it out, we're turning it off. We've got a fallback.
16:15 We're just shutting the thing down." Very sensible. Glad they did that. And in addition, they were going to put some improved testing around this particular component to ensure that if a subsequent issue came up again in the future, they'd catch this before it went out to production. We move over then to the use case site, and context here when we're considering mitigations is always important.
16:39 So, we thought, "Well, what can we do to strengthen this Caravan component to stop it failing in the way that's failing?" Well, we could go in, we could try and fix the code and solve this problem. But this was a legacy application that we were going to retire. It also represented only a fraction of the revenue. Anyway, this was not a critical business import to us.
16:58 And so, there was no real interest in actually fixing it. There was no real interest in strengthening the component itself because it didn't make sense from a business point of view. What we did at least do is put some improved monitoring around it to try and pick up these issues if they occurred in the future. What I found out about years later, but I wish I'd known at the time, was we could have put some load shedding into this component to help improve reduce the hazard.
17:24 The hazard here of Sauron as a whole was that request count. We the resource saturation. The resource saturation was caused by having too many requests. Well, we could have just shed the load. So, with load shedding, you basically say, "I can only handle 30, 60, 50, whatever, how many requests you can handle. Anything over and above that, I'm just going to drop on the floor."
17:46 And that would have kept our system up and running. So, if we'd applied that kind of load shedding protection to the Sauron application, Sauron would have carried on working. So, it's always nice. Wish I had a time machine right about now. Learning about going back in time in the previous talk on this track. You should watch that online if you haven't already.
18:06 So, how can we strengthen We've reduced our hazards, so how can we strengthen our components? How can we reduce the chances of those components failing in their entirety? And strengthening things is a good idea. We strengthen concrete with, you know, steel in it. That makes it stronger. And this was one of the major problems with Ronan Point. When it was built, there were not enough people with the skills required to build buildings how we normally build buildings.
18:32 So, a system was selected that would allow lower-skilled people to build high-rise buildings quickly. I'll say that again. We came up with a system that allows lower-skilled people to build high-rise buildings quickly. What could go wrong? So, in the immediate aftermath of this, it's like, maybe we should look at whether or not we should be doing this, full stop.
18:57 Um because it turned out that there were some significant issues in how this building had been put together. These prefabricated sections were brought onto site, and they were sort of slotted into each other like Lego. You had these ties, these steel ties that would overlap, and then the joints between the panels were supposed to be filled with concrete.
19:18 Yes, concrete has very particular material properties. Now, the lower-skilled people that hadn't had very good training, and where there was no real oversight, had made some different decisions on the fly, and thought, well, we could use concrete, or we could use newspaper and cigarette packets. I don't know if you're aware, they have different tensile properties to concrete.
19:40 This was only found, by the way, because there was an almost forensic deconstruction of the building later on that discovered this issue. This, as a result, strengthened and improved building regulations in the UK. Um had a lot more oversight. They also started instituting things like tool chest talks, almost like stand-ups at the beginning of the working day for construction, recognizing that whether or not the large panel system itself was flawed, it was not implemented the way it should have been.
20:08 So, this actually changed the building industry in the UK and beyond, because we all got to learn from this. So, how do we strengthen a service? Well, one way we can strengthen a service is by improving the redundancy in the components that we use to operate that service. We're aware, for example, that failure of a service instance is a thing, and so we decide to have multiple copies of a service instance and stick it behind a load balancer, because the idea is that if one of those instances fails, our service can
20:39 continue to operate. The service as a whole has been strengthened through the provisioning of redundant resources. This is so much an obvious known issue. Great. So, let's come back to caravans, for example. Could we have strengthened our components there by running redundant copies? Well, yeah, absolutely. The caravan site was actually a read-only application.
21:01 Running multiple nodes would absolutely have made sense. Do we want to spend money running redundant nodes of a service that we're planning to retire and that represents a fraction of our traffic and our revenue? And no. So, financially this was absolutely a non-starter at the time. So, there were some other things we're going to have to do here. What about in AWS?
21:23 Strengthening the component. Well, actually there was some redundancy already within the existing DNS solution. Those multiple DNS enactors was a source of redundancy. Weirdly, of course, that actually also caused the race condition. So, strengthening the component in the context of AWS specifically would maybe involve fixing that race condition. But then we have to look broader.
21:45 All those people that are impacted by their use of AWS. If you think about the cloud as the component, how do I strengthen my cloud? If I am the, you know, the CEO of Sleep8 or whoever the smart litter boxes, what else could I do? In the aftermath of the region failure, everyone started saying, "Oh, you need multiple sites. We need to be running across multiple clouds."
22:10 The companies themselves were a bit quiet about whether or not they were going to do this, but everybody on Hacker News thought this was a good idea, so clearly it must be a good idea. And at first glance, this does seem to make sense, right? If you look at running on AWS, we have the concept of multiple cloud regions. This giving you kind of sort of not only you know, being able to operate data in different countries, but just separating out your blast radius.
22:34 So, we could, why not run our system in multiple sites, in cross multiple regions? Well, even when you decide you're going to do that, you've got to decide what technique are we going to use? There's a whole plethora of different ways that you can run your system in a multi-site setup. This is traditional DR type stuff, right? But we can start with a classic backup and restore mechanism.
22:54 We could back up our data from US East 1, and if something goes wrong, we can reconstitute that somewhere else. Um guess what? That takes ages. And also, if it's going to work, you have to have invested a lot in automating your infrastructure. Plus, you also need to have know you're going to get the infrastructure somewhere else. But, most people should be doing what everyone should be doing backups anyway, so this is typically a good starting point if you're looking at multi-site solution.
23:21 We've got things at the pilot light set up. What the pilot light set up is trying to do is reduce your window of data loss. If you back up data once every hour, and you reconstitute your system from those backups, you can lose up to an hour worth of data. With a pilot light set up, you're more constantly replicating data from your primary site to your failover site, but you don't actually have any infrastructure running to serve your traffic.
23:45 The idea being that you shrink that window of data loss. The idea then is when you have a failover mode, you can spin up the infrastructure. It's about keeping your costs low in a dynamic provisioned environment. You then got variations, the warm failover, you know, you have some infrastructure provisioned so that you can you can sort of failover some traffic or some critical paths.
24:07 Um, again, it's all about balancing cost. But, of course, the holy grail of this is, well, let's just go multi-site, Sam. Let's have two sites running and allow traffic to be sent to either site because that's going to eliminate my data loss window. It never eliminates it entirely, but also, I'm not going to have any downtime, assuming one of the sites has enough capacity to handle all of the load.
24:30 Of course, this looks great. The problem here is the complexity. I need to be really clear about this in case you've had conversations with any vendors this week. If anybody tells you that two-way data synchronization is really easy or simple, they are trying to sell you something. Unless you started off building your application with this in mind, retrofitting this into an existing system is not trivial.
25:01 Okay. So, do your research before you think about whether or not this is for you. Fundamentally, when we start looking at these multi-region, multi-site options, as we come up these patterns, we are reducing our downtime, we're reducing our data loss windows, but we are increasing cost and complexity. So, as we go from left to right, we're getting better from a point of view of downtime and data loss, but we are getting worse from a point of view of cost and complexity.
25:30 So, a lot of the time, this comes down to money. Why weren't more of those big name companies running a multi-site setup? Money. In part. Those smart people at these companies, do you think they didn't know that a cloud that they were based in one region? Of course they knew. They absolutely knew. And they had made a decision not to have a hot failover.
26:02 There were reasons behind that. Some of that might be down to cost. Are they going to charge you more for the work required to do this, when these failures occur very rarely? Maybe not. There's other elements though to doing multi-site that we'll come back to a little bit later on. By the way, if you want to do I got a lot of good guidance around some of these patterns from this really nice write-up at the well-architected framework AWS.
26:27 These patterns are fairly generic. They don't obviously talk about multi-cloud. Weirdly that AWS not talking about multi-cloud, but there is some good stuff in here. If for some of you who went to Martin's keynote yesterday, you might also be thinking about data sovereignty. If your concern is around wanting to run multi-region or multi-provider for data sovereignty reasons, that is completely legitimate, but But, not what we're talking about today.
26:52 We will come back to that concept though a little bit later on. So, let's look at our third way of mitigating our problems, which is reducing interconnection. And if we think about, you know, that little nice little conveyor belt of dominoes, if you want to stop all the dominoes getting knocked down, you can just go and remove one of the dominoes, and you get a bit of a collapse, but it stops.
27:15 So, reducing the interconnection in our system can be a great way to make sure that a problem in one area doesn't cascade through the entire system. Within Ronan Point, the problem was that there weren't alternative load-bearing paths. The load bearing was done on the outside wall. Once that wall is removed, the rest of the floors came crashing down.
27:36 So, by establishing alternative paths through which the load could travel, you avoid that problem. So, a load-bearing wall being removed isn't going to result in a catastrophic incident. In our IT systems, in our distributed systems, we'll talk a lot about bulkheads. Bulkheads is a metaphor that comes from shipping. This is This picture is of a submarine.
27:57 So, if you hit a rock, water starts pouring in, you can close that compartment, and assuming you're not in the compartment, you know, if you are in the compartment, well, bad news. But, the rest The idea is that the ship remains seaworthy. Famously, at least there was at least one theory that the reason that the Titanic sank was that its bulkhead design was actually quite seriously flawed.
28:17 But, the idea here is that we compartmentalize the failure. We reduce the interconnection. So, if you're the owner of a smart sandbox, how do we introduce or decrease our interconnection? How do we reduce the interconnection? Well, how about having either a local fallback, or even better, being local first? The CEO of the smart mattress company, I think they're called Sleep 8, said, "We are going to have a local fallback."
28:50 Local fallback's fine, right? This is effectively a a bulkhead that closes when there's an issue, which is what bulkheads do. Local first might be even better than that, but that's a start. Okay, so what about multi-cloud? Because with multi-cloud, my my concern is what if a cloud vendor fails? The challenge always when we start looking at multi-cloud around reducing, cuz that's me reducing my connection to a particular vendor.
29:18 The problem with a multi-cloud approach is you're taking all that complexity that we've already looked at in terms of multi-site, but now we're overlaying the need for you to have skills and expertise in multiple cloud vendors. Or for you to curate effectively a commoditized layer. I know Martin yesterday talked about this idea that Kubernetes is effectively a commodity layer for compute, and he's sort of right.
29:43 But we also know that's not the whole story. You can't just pick an app and an application here and stick it on any other Kubernetes cluster. There's a lot more stuff that has to go in. That's investment you have to put into. Beyond the fact that operating EKS is quite different to operating AKS. So all of this gets worse if you throw multi-cloud into it.
30:02 So it's If you're interested in looking at multi-cloud, it's really important that you see this through the lens of hedging your risk against a supplier failing, not a site failing. And this might be the risk you're in. You're reducing your interconnection on a particular vendor. And again, there could be great reasons to do this around things like technical sovereignty or your concerns around the Patriot Act or the Cloud Act, for example.
30:29 There's this annoying problem, though, that even with multi-site setups or multi-cloud setups, that they're in they actually we can cause problems around interconnection. We are replicating data across those sites. That can actually cause problems. There's a reason why there are very few cross-regional cloud-based services provided by any of the big cloud vendors because they recognize that any connection between sites becomes a protect protect a potential sorry infection vector.
31:01 But let me simplify that down in a way. If you deploy the same cloud the same code in vendor A as you do in vendor B, your code itself is a form of interconnection. A bug that takes down your system on AWS could just as likely take down your system on Azure or your own on-prem system. This is why some firms go even further to reduce their interconnection.
31:25 We can look at someone like Monzo. A fully you know, for those of you who don't know, they have a banking license, right? They actually have to be certified that they go have the regulator come in and approve what they're doing. One of the things they have to provide is continuity of services. So, rather than taking the services of a British bank and running it in a US cloud region and thinking that's me done then, they've gone a bit further.
31:48 They've said, "Okay, yes, we're a bit worried about having an issue with a vendor or a site going down, but we're also just as worried about a bug in our code causing a problem in other people's systems. So, they've created a standing system that is a clean room re-implementation of the critical banking functionality. So, no shared code, very different architecture, stripped down, very simple, that runs in GCP with a core system runs in AWS.
32:20 The idea is and they can constantly validate their standing. So, when you go and use Monzo, you might be opted in to using the new site. What about with our with our used car site? Well, we started doing things like separating our connection pools. So, we had multiple connection pools for each downstream server. We also started using circuit breakers.
32:43 So, many of you know about circuit breakers, but the idea is is that after certain number of calls, we blow that circuit breaker open and make sure the traffic doesn't happen. We You can think of a circuit breaker as a way of triggering a bulkhead on demand. When that circuit breaker is open, we've reached a threshold, we can now move back to failing fast.
33:04 I will say there are a bunch of issues around circuit breakers. They aren't simple things. They can cause more problems in your system. I'd fully recommend if you're using circuit breakers to read these two posts by Mark Brooker. He points out some of the challenges around them and actually recommends some simple mitigations like only using circuit breakers around retries or else considering using things like token buckets, which can be much more elegant ways of handling this.
33:31 Actually, adding circuit breakers can make other problems worse. And this is an inherent paradox we have in our systems. As we increase the system's complexity to handle issues, we can end up introducing more sources of failure. And this is a real problem. We think we're worried about a service misbehaving. We have a circuit breaker that opens when that service misbehaves.
33:53 Fantastic. That's great. But then we start seeing problems like it makes partial failure worse. If I've got a service endpoint and some of the functionality works really well and some of the functionality is erroring, the functionality that's erroring could cause a circuit breaker to open. And now the functionality that was working can't be used. We also have that kind of general boom and bust problem.
34:19 All of these clients start triggering their circuit breakers at once. Our service goes from being overloaded to not loaded at all. And we can get those floods coming through. Circuit breakers being incorrectly configured can cause retry storms, which can in turn bring systems down. We think, you know what? I want to make sure my service doesn't fail.
34:41 The way I'm going to handle this, Sam, is by I'm going to having multiple copies of my service. I'm going to do that with Kubernetes. Well, what does that look like? Hello. This is your system on drugs, right? This is the stacks that we now run, and every single layer of these stacks makes sense in and of themselves. But as we strive to make our systems more resilient, we can actually end up increasing the complexity of our applications to the point where we introduce new sources of problems, new sources of errors.
35:18 This is something that David Woods has talked about a lot. This is to an extent not solvable. We just now have to be aware of it, and sometimes the right decision is to say, "That's not for us." Yes, we can do it, but we're going to make a conscious decision not to. It is completely okay for you to make a conscious decision that multi-cloud is not right for you because you're worried about the ongoing complexity you're going to have to manage, or that maybe the increased interconnection of those systems might cause
35:48 more issues down the line. You think about the AWS outage. They decided to run multiple copies of that DNS in actor. Why? Because they knew that availability zones could fail. So, they ran multiple copies. If they only ran one DNS in actor, this wouldn't have failed in this particular scenario. There wouldn't have been a race condition. It wouldn't have been possible for that race condition to exist because plans would have ended up being executed sequentially.
36:21 So, that complexity they added allowed the race condition to occur, which ended up bringing that system down. It's really annoying, isn't it, when you get into these sorts of things? And so, my kind of urging for you all of you is is not to say that you shouldn't do this stuff, but that you make it a conscious decision. There's a lot of learning by rote that goes on.
36:45 We're often time poor. We've got a lot of stuff going on. So, we just do things because other people have done them or we saw them in a conference talk. We don't necessarily understand the realities of what that causes. Finding time every now and then to have a chat as a team and say, "Is this the right thing for us to do?" is a good idea. And no, sticking circuit breakers everywhere isn't necessarily going to solve your problems.
37:09 Might give you some new ones. So, I haven't tried to kind of boil the ocean with this talk. There's a huge amount more we can talk about in terms of how we can mitigate uh concrete practices to mitigate progressive collapse. The reason I did this talk is I always find it interesting to look at the intersections between things, to look at different industries, different domains, and see what we can learn from them.
37:36 Metaphors like the bulkhead and circuit breakers, these are concepts that come from industries not like ours, but they are now key parts of how we think about our own system resiliency. So, if we're thinking about progressive collapse, progressive collapse is a concept that comes from the building industry, but there's still things that we can learn from it.
37:57 These progressive collapses seem innocent when they first start, but the results can be significant. You've probably all experienced something like this. That small failure which results in a collapse in the wider system. It's not a fun time. So, what can we do if we're going to mitigate a progressive collapse? Number one, reduce hazards. Stopping things breaking is still a good use of your time.
38:27 It's not by itself sufficient. Two, strengthen the components. Make them more resilient when those hazards still occur. Number three, we can reduce the interconnection between things. We've looked at things like reducing hazards, you know, yeah, that load shedding on sound would have reduced the hazard. It would reduce the load on that component and stopped it from being brought down.
38:53 We've looked at the ways to strengthen components, things like having multi-site, having redundancy around individual parts of the system to strengthen up those individual components. And then we looked at reducing interconnection between our systems, but it's a bulkhead, circuit breakers, multi-clouds. Uh in this I didn't go into the detail of it, but in the sound application we actually had ring-fenced bulkheads effectively around the connection pools for each of our downstream services.
39:22 What happened at Ronan Point was a tragedy. Four people died. It was honestly a miracle it wasn't more people than that when you see the scale of that. I mean, the real tragedy though would be not to learn from these things. The only reason I was able to do this talk is because there was an inquiry. There were reports done into what happened. Journalists and pushed to find out what caused this.
39:48 They wanted to try and move people back into this building after this occurred. It took a lot of fighting to convince them that the issues were endemic. It wasn't a gas explosion the problem, it was the fundamental construction of the building. It was only because an architect pushed to have the building brought down and broken up piece by piece that we found out all the things that had gone wrong cuz a normal demolition would have just involved blowing the thing up.
40:12 >> [clears throat] >> I only got to talk to you about the AWS US East 1 outage because AWS put out their incident report. You can see incident reports as PR and of course that's part of it, right? There is a bit of part of that. But this is also how us, we as an industry learn from each other. I learned so much from reading these incident reports. They're always fun things to read cuz also like it's not you, so that feels good, right?
40:37 That's always nice. The Cloudflare ones are especially good cuz they're really self-flagellating Cloudflare. They beat themselves up, probably a bit too much. Like if they were a person, I'd go and say, "You're all right, hon." cuz it seems like they're always good to read cuz we can learn things. We can live vicariously through them. And all these other reports and these articles and videos, they all gave me an opportunity to learn about what progressive collapse was and apply these concepts in the my own way and how
41:01 I think and how I operate. You're never going to eliminate failure from your systems entirely. You're going to do your best to manage it. Bad things will happen. But it's really important that you create the opportunity to learn from that and to also share that learning with your colleagues and if you can, the wider industry. And that's why I started doing conference talks.
41:23 It would've been 20 years ago now. If you want to know a bit more about Rowan Point, this video is a great watch. I'm going to give you a slide download at the end. I wish I'd found that video before I went through all the PDFs cuz it does explain a lot of what happened in a bit more detail in a nice 12-minute video. Um uh but these and other resources are available as links from this slide deck.
41:46 You can download the slide deck here, but it's also available to download via the QCon platform. I'd really appreciate it if at the end of the talk you can give me some feedback. The things you liked, the things you didn't like. I'm more interested in what's in the text box than I am the thumbs. It's a personal thing, but you do you. Um my new book will be out if I wasn't here, I'd be writing it.
42:06 Uh should be coming out in uh Q3 this year. Although about 2/3 of the Well, about half of the book is already available to read if you're on the O'Reilly platform. But, I hope you enjoy the rest of your time at QCon. Thank you. >> [applause] >> Thanks very much. Really interesting talk. Um very interesting to see you talk about how adding capacity to avoid issues can, you know, create issues in itself.
42:32 I'm also wondering what your thoughts are on the time scale to kind of like think about adding protections because when you have an incident and you have a, you know, a retrospective, everybody's inclination is to be like, "Oh, we need to make sure this never happens again like right now." And I'm wondering like some of the things you talk about like uh you know, have longer time scales or you need to think about them more.
42:55 I'm kind of curious like how you approach advocating for certain kinds of protections um with that in mind. >> I I I don't think in the wake of an incident going, "Let's make sure this never happens again." is a bad thing. Right? I I've said this before, but closing the stable door after the horse has bolted is always a good thing. Ideally, we would have closed it before the horse bolted, right?
43:22 But, we can't go back in time. So, yes, maybe we should close the stable door. And a lot of the time it's like, "Oh, I've got a stable and a horse, right?" So, I think there's I I that's understandable. But, I think it's more about prioritization. It's more about like when you increase the complexity of a system, that has an implementation cost that has an ongoing kind of ownership cost.
43:43 So, I think it's like any other piece of work. What is it we're trying to fix here? Getting concrete can help. I didn't go into it in detail, but when sort of doing DR planning, for example, it's useful to separate out like your restore point objective from your um you know, which is when you've kind of last week your back up of data. So, how much data loss can you can you qualify?
44:03 And And going to be a business requirement, right? How what's my data What's my window of data loss? Or my restore time objective, which is how quickly do we need to slide up and running? A lot of the time those kind of um operational requirements are things which you can go and find out about. And then you can use that to prioritize the work. I think as techies we are often thinking about this world.
44:29 And I think it's if we if we put the the context of the failure in a business cons concept or business context get a bit more specific about what our system needs to be able to do and then break down the work and explain how we're going to achieve it and then it's prioritization. This thing is going to help us make a bigger impact than this thing and it's going to be less work.
44:48 Let's do that first. And then for me it just becomes a conversation with the product owner or business owner about you don't like that, do you? This will fix it, but it's going to cost you 5 million quid. I haven't got 5 million quid. Okay, we're agreeing not to do it. Let's move on with our lives. I just think it's that's that's simple. It's not simple, but you get the idea.
45:10 I think any other questions? We got number two. Hello number two. >> Hello. Uh thank you for for the talk. Um there was one sentence that uh you posted that um I haven't been able to strip out of my out of my head, which is at some point in the building issue uh there was uh the ability for lesser skilled people to build uh the building uh faster. >> Yep.
45:35 >> And it kind of makes me think about the current situation where with the use of agents and AI we are able to get less skilled people to build solutions software solutions uh faster. >> Well, we're doing we're doing that by not by by getting rid of junior developers and not hiring them anymore. We're doing that from two different angles. It's great, isn't it?
45:54 Like >> So, I was wondering in this context if you have any thought that could be applied in the current situation. Uh um what guardrails we can put in in the case of agent building software. >> Well, that's coming back to the analogy of the construction, right? The The The two sides to look at what I'm pointing is was the system that they chosen for building the building correct in the first place?
46:21 And the second thing is did they implement the building based on that system? And there's question marks about both things. And I think we can say the same thing from an AI point of view. So, the models themselves are getting better. Right? So, hopefully what they're producing is better. We're getting a better understanding about how to guide it from a context point of view.
46:40 So, the system itself we're hoping is a better system. We still need the guardrails on the other end, the verification. I was at a panel yesterday from from the cohort training. And like I actually I think there's a bit of agreement amongst us on the panel is that we've we need QAs again. Right? The whole shift towards the whole move towards shift left around testing had this problem that we suddenly said, "Oh, we don't need QAs cuz developers are going to do all that."
47:10 No, proper QAs who can think about system properties and verify is the system working correctly. Like if I had advice for a junior developer right now who can't get a job, it's like become a QA for this world. Right? A technical QA that can verify the outputs of these systems or become an SRE. Becoming an SRE is very difficult and you've already got a job.
47:31 So, I think we need that. I think we need those skills and disciplines back. And we need to find new ways of understanding and verifying our software and making it correct. If you want to run an agent swarm that's going to run 2 weeks and rebuild a bad browser, you know, that's a well-defined problem space. If we try to do that for our own software, we need to get a lot better about saying what good looks like. I don't think we're nowhere near it yet. We're getting there. >> [music] >> Mhm. >> [music]