← All transcripts

Every AI Company Is Accidentally Building a Bank — Dor Sasson, Stigg Transcript, AI Summary & Key Points

AI Engineer · 2 days ago · Science & Technology · 20:08 · EN

Watch on YouTube

Answer

Every AI company is accidentally building a bank in the sense that AI monetization now requires banking constructs: synchronous entitlement checks before inference, asynchronous settlement, hold-and-settle for concurrency, double-entry bookkeeping, credit pools with drawdown priority, and hierarchical budgets and spend caps.

AI Summary

Every AI company is accidentally building a bank. In April, a wave of pricing emergencies hit AI products — Anthropic cut third-party agents like OpenClaw off its subscription plans, OpenAI bumped its Pro pricing 5x overnight, and GitHub froze then eliminated free Copilot access — because checks on what users and agents are entitled to happen after the invoice, not before. This is really an infrastructure problem, not a commercial one. The fix is the banking model: check entitlements synchronously before inference, settle asynchronously after, like an ATM that verifies funds before dispensing cash. OpenAI's own financial engineering team published a similar architecture in February, with a runtime decision waterfall. AI products now need banking constructs — hold-and-settle for concurrency and double spend, double-entry bookkeeping, credit pools with different sources, hierarchies with budgets and spend caps across teams, users and agents. Four patterns repeat when selling AI at scale: reserve before inference and settle actuals after, keeping the ledger in your VPC, checking balances continuously for agentic work, and spend visibility that is now table stakes for CFOs and CIOs. The industry is missing its Stripe moment for financial infrastructure.

Key Points

  • In April, within about 5 weeks, every code-generating product platform broke under the new AI economy: Anthropic eliminated third-party agent access (including OpenClaw) to subscription plans, OpenAI bumped pricing overnight 5x, and GitHub froze then removed free Copilot access.
  • The common root cause: entitlement checks happen after the invoice, not before — users and agents consume first, and spend is reconciled only after it is too late, causing sticker shock, overspends and margin problems.
  • Anthropic/OpenClaw case: customers paid dollars and cents daily while Anthropic subsidized $150 to $750 per user, and Anthropic lacked a way to distinguish API users from subscription users, forcing it to cut access — a business problem that was really an infrastructure problem.
  • Other examples of unchecked AI spend: an undisclosed company burned through half a billion dollars of AI cloud credits with a few employees, Uber torched its entire yearly AI usage budget in a few weeks, and three Replit users in one org burned down the entire organizational credit pool.
  • The correct architecture: enforce checks synchronously before updating balances, and reconcile and settle asynchronously after — like an ATM that checks whether you can draw cash before dispensing it.
  • OpenAI's financial engineering team published this architecture in February: a runtime decision waterfall evaluating feature access, rate limits, trials and promotions before allowing drawdown — single synchronous evaluation, deterministic priority, and not merely billing but a financial system.
  • Concurrency and double spend: banks solve simultaneous drawdowns with hold and settle, while 10,000 agents hitting the same credit pool simultaneously is a classic double-spend problem requiring double-entry bookkeeping, idempotency and auditability.
  • Credit pools are not fungible: a pool of 1,700 credits may come from different sources (grants, promotions), and companies must decide drawdown order — most today store it as a single integer in a database, which will not scale.

Business ideas

Every AI company is effectively building a bank: credits, holds, settlement, ledgering. The core architectural change is moving entitlement checks from after the invoice to a single synchronous evaluation at request time — like an ATM checking balance before dispensing cash — while reconciling actual usage asynchronously afterward. The speaker argues the industry misses its 'Stripe moment': an easy-to-adopt set of banking constructs (hold-and-settle, double-entry bookkeeping, idempotency, auditability, credit pools with drawdown ordering) that any AI product can plug in instead of meeting these problems at scale, when they are far harder to fix.

For
AI product teams selling usage-based or credit-based AI workloads, from startups to enterprises whose buyers (CFOs/CIOs) demand spend control.
Solves
Entitlement and spend checks happen only after the invoice, so consumption is reconciled when it is too late — causing pricing emergencies, overspends, margin erosion, sticker shock, and emergency access freezes when customers burn through subsidized credit pools.
  • Anthropic and OpenClaw: Anthropic was subsidizing OpenClaw users' consumption under Claude Max subscriptions — customers paid dollars a day while Anthropic's cost ran to hundreds per subsidized user — forcing it to cut off third-party agents from subscription plans.
🔒  Build steps and tools for 1 idea. Unlock

Tools & resources

2 items

Links mentioned

🔒 Full analysis locked

Unlock more videos and the full analysis

A credit unlocks one video's full analysis for good — the build steps, the tools and how each was used, the methods behind every use case. Pro opens the whole library instead, and raises how many videos you can analyse a day.

Unlock full analysis — free

Transcript

Searchable transcript of Every AI Company Is Accidentally Building a Bank — Dor Sasson, Stigg — AI Engineer (20:08). Search for a phrase, then click its timestamp to jump straight to that moment in the video.

Captions sourced from the original video on YouTube, published by AI Engineer. The video, its captions and all related intellectual property remain the property of their respective owners; AINotes claims no ownership. Provided for research, accessibility and search — see the Transcript Notice and Copyright Policy.

00:12 Yeah, we're good. Okay. So, hi everyone. Given there's I hear there's a match right now as we speak. So, given that you are here means a lot to me. So, at least I'm I'm winning the World Cup right now. So, that's that's a that's a big thing. Thank you for joining me. I flew all the way over on co-founder CEO of Stig. I flew all the way over to make a kind of a funky statement on an AI conference which is literally every AI company right now is accidentally building a bank.

00:41 And when I mean bank, I'm going to try to kind of unpack this thing not as like a weird marketing metaphor, but effectively what what is really happening right now with the AI economy and why a lot of the constructs and the things that we are seeing using consuming and paying for in AI are actually really behaving like constructs that we know from the banking financial systems.

01:06 So, cool. Thank you for for for having me. Um and let's get it going. By the way, I'm the first talk today. So, the clicker is not working. It's going to be interesting here and out as they move for this. So, let me take you through a little bit of background. As I was working through this talk back in April really every single code generating product platform out there just broke under this new type of economy and under this new type of consumer behavior.

01:38 And in just a you know, matters like 5 weeks everything felt like pricing emergency one after the other. And the case I want to make here is that these pricing emergencies go way beyond just emergencies from a financial commercial set of the house. These really are infrastructure emergencies that are emerging because of how those systems were built to begin with.

02:02 So, just to give you some perspective, we had in in April we had Anthropic basically eliminating access for Open Clo and other third-party agents to basically use the subscriptions plans. Then immediately we had OpenAI basically changing how the product here is being priced, bumping the price overnight 5 5x, which is quite intense. And then we had GitHub freezing and then completely eliminating access trial access free access to their We have some echo here.

02:35 To their free plans for for for Copilot. And effectively like all these companies all at once in a single month experienced what it means to actually scale on scale AI and what actually happens when you don't have the fundamentals in place. So, let me take one one step farther. So, let's zoom into even just one of the use cases, right? So, if you think about Anthropic and you think about the Open Clo use case, right?

03:05 What happened was effectively Anthropic was subsidizing every consumption of every user using Open Clo within the Clo Clo Max subscription. So, effectively what the customers were paying daily was, you know, dollars and cents. But on the on the cost side for Anthropic, it was like 150s to even 750 subsidized, which is effectively not an economics that you can scale even for the company at at the size at the scale of Anthropic, right?

03:34 So, what it seemed to be a business problem really was an infrastructure problem because Anthropic didn't have a an effective way to delineate between users who are using their APIs and users who are using their subscription um, uh, programs. And so, they had to immediately cease and stop access, right? They had to to to react. Um, cool. So, and this is not unique to Anthropic, guys.

04:04 Like, uh, we're seeing this like, I don't know if you all heard read the news most recently with a company, uh, undisclosed company burning through like half a billion of their AI, uh, cloud credits, um, just merely having like a few employees burning through the entire organization contract. Uh, we had, uh, Uber saying they're basically torched down their entire, uh, AI usage consumption budget for the year in just a few weeks into the year.

04:30 And then we had Rep- Replit giving an example of what could happen if, uh, a single org has three users that basically burning down the entire credit pool for Replit for the entire org. So, really what happening here is, uh, this is not just an an an Anthropic problem. And I think what's what's mutual and kind of like switching gears from the examples to the actual what's really happening here, um, the problem basically with with all these three examples is that the checks for what you're entitled, what you're allowed

05:05 to do, whether you're an agent or a user, we are AI product happened, uh, after, uh, the invoice and not before. So, basically with each and every one of those examples, you were allowed to consume, you were allowed to use the product, you were allowed to burn tokens, and checks happen only after the invoice. So, basically the spend and the aftereffect gets reconciled only after it's too late, only after the invoice is is sent, and this is effectively bad for business, it's sticker shock, it's bad for practice of how

05:35 we use these systems. These systems are now mission critical, and we're basically none of these companies had in place the financial infrastructure that allows, exactly like banks, guys, to check before we draw down, before we consume, right? So, I think what what we're seeing is effectively uh really is a change in how we architect, and how we ship, and how we build software.

05:59 AI is not just changing what is value, is not just changing how we interact, who is the users. It also effectively changing what exactly is getting paid, and how we actually architect for systems that can uh handle runtime, handle the complexity of usage happening before the the the invoice, before the checks, and how we actually can architect for that.

06:25 So, I think what's interesting, and what you see in front of you um is basically on the left side, what we're seeing is how most teams ship this architecture today, and you can see that basically the check and the settlements are happening after after the inference, after the usage is being logged. And that's typically too late. We see situations like spikes, overspends, overages, like experiences that are bad for business, bad for your users, and quite frankly, in many cases, also bad for the vendor.

06:55 Because if you didn't allocate enough resources and compute to actually allow for this to happen, you you also have a margin problem potentially, right? On the right side, we're basically saying, or I'm trying to make the case for you all that with AI, with the current infrastructure, and how AI is being built and sold, we actually need to enforce and and do this synchronously before we actually update the balances, and reconcile and settle in in async way after after effects.

07:26 So, hot path actually get checked synchronously, and everything that comes after uh is basically reconciled after. I think the best example, guys, is if imagine you go to the ATM, right? And you draw cash. When you draw the cash, it that the check if you are allowed to draw it, doesn't happen after you get the cash. It happens before you draw the cash.

07:46 And so, that's exactly the same thing. Our industry effectively misses this layer, this financial infrastructure that allows for this type of behavior to happen naturally, right? And so, obviously this sounds like a great idea, though, right? Like you're you're pitching this infrastructure, sounds like you have something to do with this. Um But guys, this is not just me.

08:09 Uh in February, OpenAI, they have a quite a great team. They're team called financial engineering. And these guys basically published how they think about the architecture and what they had to build for this to work at the scale of OpenAI. And what they had to say reads very similarly to how transactional systems of banks work. Basically, you have a decision waterfall that needs to consider in runtime what is the client, the user, the agent is allowed to do.

08:42 And it needs to take into account many different conditions and policies of what is really allowed, right? So, are you allowed or you have access to a certain feature or product? What are the certain rate limitations under your plan? Whether you are in a trial or in a certain promotional program. Like you need all these decisions to be calculated immediately at front time.

09:07 And then it ultimately decide, can I give this user access? Can they draw down? Can they consume or not? And all those decisions really can happen after the invoice. They can't because you need to actually be able to calculate and prioritize this. So, I think we have three things to notice from here, right? The first thing is this is a single synchronous evaluation.

09:26 It happens at the request. It doesn't happen after effect. I think the second thing is the priority is deterministic. So, we know in advance what are all the different rules that apply for certain things to be consumed, right? And I think lastly, and here you guys, this is my opinionated view, you can disagree. I don't think this is billing. I don't think this is necessarily this entire architecture is just related to the idea of invoicing and billing.

09:56 This is a financial system that actually make decisions ahead of any invoice, ahead of any billing element. And I think we're going to see more and more of this becoming a standard for other products and other AI systems as well. Really just nerding just a little bit on I'm sorry, I just skipped one slide. So, yeah, just just nerding just a little bit about what other trades or or elements we see and are used to in banks and we take them for granted, but in AI they're not that granted, in software they're not that

10:30 granted. So, just one example, right? Like if two separate parties have access to the same banking account, and that banking account has like $10, and they are entitled to draw down those $10. What happens and when they both of them try to draw down at same time, right? So, banking systems has the idea of holding, of settling. You can really the concurrency is effectively solved with banks, but with AI and agents, it's not that case.

10:58 You could have whatever, like 10,000 agents trying to reach out to the same dollar pool, and they're going to try to do that all at once, and superficially it might look like they should be able to draw down because there is balance in the pool. But, what happens when they all draw down from the same pool simultaneously, concurrency become an issue.

11:16 So, how do you solve for that? This entire idea in financial systems actually called like it's a classic double double spend problem. It comes from the accounting world. It's a double entry bookkeeping, and as I'm sure you all are building shipping today with AI as the revenue scales, as your company scale, the idea of double entry accounting, double entry billing, the idea that you can solve concurrency, these things would matter as you scale and they're not so easily solved down the line.

11:47 So, think about things like you know, can you basically introduce rates like hold and settle, think about idempotency and auditability of those uh requests and etc. Another thing that I think is very very known in in you know, how banking systems works but not not necessarily enough established in AI is the idea that all the different pools have different sources, right?

12:18 So, in bank you have your debit account, you have your cash account, you have your savings, right? And if you draw down, the bank actually knows how to calculate where you drawing down from and what does it mean? Here, you know, say I have a pool of you know, one you know, 1,700 credits but they're not the same, they're not alike. Each one of them is sourced from a different source.

12:43 So, how do you decide which which one of them do you draw down against? Do you start with the ones that were you know, given granted earlier? Do you start with the ones that were promotional? So, how do you really calculate what's getting drawn down to the users? I think that's that's a real problem. Most companies today try to solve it like a single integer databases and that's not really going to work as you think forward.

13:08 Cool. Um so really tying the knot the knot here, I think if you think about how we interact with software, it used to be flat. So, it used to be just users and organizations and there wasn't a lot of meat in between them, right? The the the hierarchy of you consume software was quite flat. It was quite straightforward in terms of who see the value. But today, I think effectively credit pools are no longer flat.

13:34 Uh you have a lot of different hierarchies in between who consumes different credits, and also they demand different governance, different logics, allocations, budgets, spend caps. And effectively, what happens is that budgets are set at the top, where for instance, if you're selling to the enterprise, you're maybe selling like pre-commit annual contract.

13:56 Uh C-suite or somebody who approved that seven-digits contract, they expect to see how that credit are getting drawdown along the contract life. But they don't just want to see that they were used. They want to know who actually used them, which team, which users, uh which agents, uh to what extent, and they want to have fine-grain control over that.

14:17 That's literally the bare bones of how more and more companies thinking about selling AI workloads and how they they effectively work in production. I'm going to run through some of this. Um I think this idea of like what does it mean like high cardinality graphs and dimensions, being able to slice and dice against model types, users, like all this type of uh dimensionality of uh of uh consumption is is probably interesting to to talk about.

14:44 But what I actually want to expand on is ultimately those four patterns. So as you build AI and start selling AI at scale, you're effectively like four patterns that keep repeating themself, and they are very very much alike and and very similar to the idea of, you know, banking systems. I think the first one we talked about earlier is basically being able to reserve before inference and settle actuals after.

15:07 Um this is like going to to be like a pattern that again and again we're seeing more and more often, and it has ramifications on how you do accounting, how you do uh basically handle all those uh different concurrencies workloads and what have you. I think the second thing is we're seeing more and more AI companies unexcited about the idea of sending all this usage data to to somewhere else over the cloud.

15:31 For latency reasons, for cost reasons, for data sovereignty reasons, there's a lot of good reasons for why these companies actually rather to keep the events data inside their VPC, and I think we're going to see more and more AI companies seriously considering deploying this this architecture, this ledger, this metering within their VPC rather the the contrary of sending it over the internet.

15:52 I think for agentic checks, so think about how agents spawn, the idea of checking balances when agent does something you don't know in advance what they're ultimately going to do, how that how that task is going to basically cost. So, if you don't know the cost, you need to be able to check those balances as the agent continues to get work. And you need to be able to reserve an account for them in advance and then asynchronously after effect reconcile.

16:25 Quickly moving forward, so ultimately if you have those systems in place, what you really gain out of it is visibility. So, it used to be given that for usage-based products you need visibility, but I think what happens today, you know, it goes beyond just plain visibility. It it means that effectively like the ability to see for CFO, for CIO consumption of AI workloads across different models, across different features, across different products is not just, you know, a nice to have.

16:54 It's becoming a table stakes. It becomes something that you can't do business without, and we're going to see more and more companies expect to have this off the bat before they actually do business with you. And so really just giving some examples to how some of the best frontier AI companies are approaching some of this, I think you're you're seeing the patterns already there.

17:17 So, I think a lot of these companies already appreciate that you're no longer just selling software in the you know, commodity way that we used to think about subscriptions and usage. It's becoming more and more of a financial systems and financial systems require some of these complex ideas that weren't we're not there before AI. Um and I think just to give like one small example that I talked about earlier, but I think might have gone a little bit uh hidden.

17:45 Um One of the changes in April that OpenAI did is really they didn't change the pricing. All they did was change the pace in which tokens get burned. So, effectively you use the same model, we pay the same cost uh per unit, but we burn more tokens. So, that effectively was a pricing change. How do you execute such a change? Um because it's not just financial change, it's an infrastructure change.

18:09 You're basically change the rate of how the software gets consumed. So, there's a lot of complexity not in terms of just how we change pricing, but also how the infrastructure underneath allows for more complex ideas into how value is getting uh uh introduced to the users. And really just, you know, before wrapping up, you know, there our team likes to say, and this is something we started to say more and more often than before, it really feels like you know, there was in the in the early 2010s there was this Stripe

18:37 moment where every company wanted to ship product into the internet and you need a checkout and you need it to be secure and it and the developers need to to to to to make it easy for them to quickly get them up and running. I feel like where we are today is those ideas of like building a banks, it's okay if you feel like, I don't know, you like wearing suit and tie, maybe you want to build a bank, but if you are a building a bank, you probably want to be aware to all of these traits and elements that has to be there

19:05 so you don't meet them down the line where where it's get really, really more complex to fix in a hindsight. And so, feels like the industry really misses its Stripe moment where all these constructs just exist and it's they're so easy to use. And you know, there's a lot of excitement from our team when it comes to the potential of having such systems ready and help AI companies keep shipping.

19:28 Um if any of that was interesting, really come by our booth. We are in a quite of a weird location right at the start, but still there. I know there's a game right now. Come visit. We're demoing the product. Feel free to ask any questions and thank you for coming today. It was awesome to to have you have you with me. And yeah, love to nerd some more by our booth.