← All transcripts

Why smarter AI models could drive up compute prices 10x Transcript, AI Summary & Key Points

Dwarkesh Patel · 4 hours ago · Science & Technology · 11:18 · EN

📄 Transcript

Searchable transcript of Why smarter AI models could drive up compute prices 10x — Dwarkesh Patel (11:18). Search for a phrase, then click its timestamp to jump straight to that moment in the video.

Captions sourced from the original video on YouTube, published by Dwarkesh Patel. The video, its captions and all related intellectual property remain the property of their respective owners; AINotes claims no ownership. Provided for research, accessibility and search — see the Transcript Notice and Copyright Policy.

00:00 Today I want to talk about what the  compute situation for the labs will look like over the next few years. For the last three consecutive years, Anthropic's revenue has 10x'd year over year, and  it's likely to do so again this year. They ended last year with nine billion in revenue. I think  they'll probably end this year with somewhere between one hundred billion to one hundred  and fifty billion dollars in revenue.

00:14 Now, for this trend to continue, Anthropic would  need to make one trillion dollars in revenue by the end of next year. Of course,  there's no deep reason why this has to be true. It's a very wild conclusion, and  it's ultimately a question of AI capabilities. Does AI get that useful by the end of next  year? But suppose the trend does continue. I want to think through what happens in that world.

00:33 The other big trend in AI is that lab compute only 3x's year over year. For a lab to keep 10x'ing  revenue year over year while compute only 3x's, one of the following three things needs to  happen, or some combination of the three: One, lab margins have to increase. Two, the  price of compute has to increase. Or three, the percentage of compute that labs spend on  inference rather than training has to increase.

01:01 My understanding is that basically all three  of these things are already happening. With regards to the margins, Anthropic's inference  margins reportedly went from forty percent in the middle of last year to upwards of eighty  percent now for Fable. With regards to compute, the spot prices for compute are more than forty  percent higher than they were in the February trough that we had earlier this year.

01:20 And with  regards to the share of compute that goes to training versus inference, in 2024, according to  Epoch, OpenAI was spending just a quarter of its compute on inference, and that number is likely  closer to fifty percent, if not higher, now. Now, labs would prefer not to do this final  thing of increasing the share of compute they spend on inference.

01:37 The way the labs see the  world, the whole point of inference revenue is to help convince investors to give you more money  in order to train the next bigger, better model. And if you're spending most of your compute  on inference, then you're basically declaring that AI progress has stalled and you're just now  in the business of being a cloud provider.

01:53 This is a less compelling business than building AGI,  so the labs do not want to be in this business, nor do they think they are in this world. They  think that within a year, they'll have built models that make the current ones look extremely  shitty. But they need to invest a lot of their compute — the majority of their compute  — into doing the training and experiments that are necessary to build the next model.

02:16 So that leaves only two options for how you can get out of this gap between the fact that  lab compute only increases 3x year over year, but revenue increases 10x. Either  the lab's margins have to increase so that they get the surplus, or the price  of compute has to increase so that everybody in the stack below the lab gets the surplus. It's not clear to me which world we end up in.

02:39 Do we end up in a world where we go from 80% for  some of the top models to greater than 90% margins if the lab margin effect dominates? That would  require the leading model to be so far ahead of the competition, because the nature of margins  — why they exist in a market economy — is that the thing you are serving is so much better than  what somebody else could go get and replace you on the market.

03:01 But it's just really wild for me  to consider that the margins for something like intelligence will be greater than 90% and  they don't get competed away at that level. So that leaves only one other possibility of  this escape valve between these two trends, which is that the price of compute has to  increase. As I mentioned, this is already starting to happen.

03:18 And the effect is even stronger when  you look at the tranche of compute that the frontier labs actually need to accumulate, because  they can't just go out and buy a spot instance. They need to make sure that they get enough scale  to get really good efficiency and flexibility, and also that they have the kind of compute that  lends itself to the security they need for their own weights and for their customers' information.

03:39 I think a relevant case study here is to look at the compute that Google and Anthropic are renting  from SpaceX. Google, for example, is paying nine hundred million dollars a month for a hundred  and ten thousand GPUs that are a blend of GB200s and GB300s. The price that Google is paying here  is 2x the spot price per hour for those GPUs. And that spot price itself is more than forty percent  higher than it would have been in February.

04:06 I want to emphasize a key conclusion here: as  AI models get smarter, they will be better able to monetize the same amount of compute. If a true  human-level software engineer could run on an H100 equivalent, then at today's prices for software  engineers, that H100 should rent for over 250K a year. That's over 15x the current spot price for  an H100.

04:24 And this is not even accounting for the fact that your AI can work nights and weekends. Of course, you might expect that if we had ten million extra software engineers suddenly appear  in the economy, the marginal value of a software engineer would decrease, and thus the revenue  that that H100 would be able to generate would not be 15x higher than it is right now.

04:45 But  I actually don't know if this is true. If we apply this argument to people instead of AIs, then  this would be the classic lump of labor fallacy. For example, economists generally believe that  high-skill immigration does not decrease wages in the long run because of how innovation and  specialization increase the value of labor. Maybe this labor supply shock will be so big  and so fast that we can't count on this general heuristic anymore.

05:11 But if you believe what  standard economics says, then the marginal value of labor, and thus the marginal value  of compute, should stay astonishingly high. So let's think about what changes in such a world.  One of the things that would happen is that as the top labs get better and better at monetizing  compute, and the cost of compute increases, it becomes harder for anybody else to  compete against them, because they have to bid for this resource against somebody who  is basically able to make better use of it.

05:38 Another thing that will happen — and I think this  is actually the most interesting implication of this whole thought exercise — is that if you  can train the best, most efficient model, then you'll be able to charge much higher margins than  you can today. This is the Alchian-Allen effect in economics, and what it's basically saying is  that if it costs twenty dollars an hour to rent an H100, then it would be extremely stupid to use  a weaker, less efficient model, because it's gonna burn more tokens on your expensive

05:59 compute to  get the exact same result. So labs will be able to charge a much larger premium if they can train  a model that better economizes this scarce input. Basically, if you have a model that can get the  same result by using less compute, then you've, in some sense, created more compute, and  the value of compute is gonna increase. Another thing that will happen is that a lot of  current popular applications of AI will probably get priced out.

06:25 The reason AI is relatively cheap  right now is that AI just can't do a lot of things that top humans can do. But this, at some point,  will no longer be the case. And at that point, Google or Anthropic or OpenAI will be  willing to pay more for the tokens to automate AI research than you or I will be  willing to pay to make more AI slop talk. I'm a bit worried that this kind of analysis  honestly pattern matches a lot onto the ways that people in the past have been wrong  about scarcity.

06:50 I'm thinking, for example, of the famous Simon-Ehrlich bet. Paul Ehrlich  was this famous doomer about population growth, and he made this bet that a basket of commodities  would increase in price rather than decrease in the decade preceding 1990. This is a very famous  bet because it's supposed to illustrate how Ehrlich's Malthusian worldview was wrong,  and how he did not anticipate the way in which market signals and human ingenuity can  find better ways to economize scarce inputs.

07:24 I'm guessing that the analogy to this bet is  probably wrong. Other analysis has shown that if that bet had been made in a different decade,  Ehrlich might well have won. But more generally, I think the supply of compute is much less  elastic, much less capable of absorbing large demand shocks, and much less capable of being  accommodated by using different substitutes than the extraction of different metals is.

07:47 To illustrate why I think this 3x in compute capacity year over year is hard to  budge or potentially even sustain: I don't see how any of the three elements that  constitute that 3x can be much accelerated. 1.4x of that is coming from Moore's Law. Far  from increasing it, I think it'll be a miracle if we can just keep it going for a few more  years.

08:06 1.2x is coming from building new fabs. This process is ultimately gonna be bottlenecked  up to 2030 and potentially even beyond by just building new ASML EUV machines. Dylan, when he was  on the podcast a few months ago, talked about this in great detail. And 1.8x comes from the fact that  AI is absorbing a lot of wafer allocation that was previously going to smartphones and PCs.

08:30 This is  probably gonna hit a wall by the end of next year, when at the leading edge N3 nodes at TSMC, AI  will have gone from 60% to 86%. At some point, you have just absorbed all leading-edge wafer  capacity for AI, and you can't keep increasing this number. So I don't know how we even continue  to do 3x compute scaling year over year for the next few years, much less go beyond that.

08:53 At the end of the month, I go through the time-honored tradition of closing my  books. I start by opening Mercury, which is my banking platform, to make sure that  all my transactions are properly categorized. Auto-categorization rules handle the predictable  stuff pretty well. But I'm constantly working with new contractors—tutors, researchers,  and videographers—and I'm also trying new tools.

09:14 Manually categorizing all of  these transactions would add a couple of hours of overhead every single month. So instead of going through them one by one, I have Command, Mercury's built-in AI,  take a stab at all of them at once. Command proposes a category for each transaction  and provides its rationale. I just review, fix anything that's off, and approve.

09:31 Once all this work is done in Mercury, it syncs everything with QuickBooks. And  Command's judgment calls are genuinely good. It does the obvious things, like looking  at the vendor, but it also investigates who on my team made the purchase and looks at notes and  memos to build up as much context as possible. This is just one of the ways you can use Command  to automate the back end of your business.

09:51 To learn more, go to mercury.com/command. Mercury is a fintech company, not an FDIC-insured bank. Banking services are provided  through Choice Financial Group and Column N.A., Members FDIC. AI-generated responses and  suggested actions may vary and are not guaranteed. Now, I wanna clarify that at some point in  the future, compute will get cheap again.

10:05 At some point, we'll just have robots that can  convert shores of silica sand and mines of copper into new computer chips, and then the  price of compute is basically the raw inputs and the tools required to do this processing. I'm  just talking about this current pre-singularity regime where AI compute merely 3x's year  over year, which is not enough to offset how much more valuable AI is becoming over time.

10:29 By the way, the fact that Anthropic's revenue has been 10x'ing year over year, whereas their  compute has only been 3x'ing year over year, I think illustrates how strong the economies of  scale are in the model business. And logically, this makes sense. When you train a model,  you just have to spend this one-time cost to learn all these different skills that then  get to be shared across all your users.

10:47 This is very unlike human labor, where each instance has  to be retrained from scratch. I wish we didn't live in a world with such strong economies of  scale for intelligence, because I'm worried about power concentration, but it seems we do. Okay, this was a narration of a blog post that I also released on my website at dwarkesh.com. Check  it out for other posts or to be notified when I release a post in the future. Otherwise,  I'll see you for the next full episode.

💡 Answer

Smarter AI models could drive compute prices up as they generate more value from each unit of compute, especially while compute supply grows only 3x year over year and AI demand continues to expand.

🧠 AI Summary

AI lab revenue is growing roughly 10x year over year while compute capacity grows only 3x, requiring higher margins, higher compute prices, or a larger share of compute devoted to inference. Higher margins and compute prices are already emerging, while compute supply is constrained by semiconductor scaling, fab construction, and leading-edge wafer allocation. As models become more capable, they can monetize the same compute more effectively, potentially making compute much more expensive and pricing out lower-value AI applications. Strong economies of scale in intelligence may also increase power concentration.

🔑 Key Points

  • Anthropic's revenue grew 10x year over year for three consecutive years, reached nine billion in revenue last year, and could reach 100 billion to 150 billion dollars this year.
  • Lab compute grows only 3x year over year, creating a gap that must be resolved through higher margins, higher compute prices, or more compute allocated to inference.
  • Anthropic's inference margins reportedly rose from 40% in the middle of last year to upwards of 80% now for Fable.
  • Frontier labs need large, secure, flexible compute deployments rather than relying only on spot instances.
  • More efficient models can command higher margins because they use scarce compute more economically.
  • Many current AI applications could be priced out when leading labs are willing to pay more for tokens to automate AI research.
  • Compute supply growth is constrained by Moore's Law, new-fab construction, ASML EUV machine production, and limited leading-edge wafer capacity.
  • Strong economies of scale arise because a model's one-time training cost produces skills that can be shared across users.

✅ Actionable items

  • Labs can close the revenue-to-compute growth gap by increasing margins, raising compute prices, or increasing the share of compute used for inference.
  • Frontier labs can secure large-scale compute deployments that provide efficiency, flexibility, and security for model weights and customer information.
  • Businesses can use Mercury Command to categorize transactions in bulk, review its proposed categories and rationales, correct errors, approve the results, and sync the data with QuickBooks.
  • Compute analysis can be broken into Moore's Law, new-fab construction, and the reallocation of wafer capacity from smartphones and PCs to AI.

🏗️ Business models

AI model business with economies of scale10:28

A model incurs a one-time training cost to learn skills that can be shared across users, allowing revenue to grow faster than compute capacity.

  1. Invest in training and experiments to build a capable model.
  2. Share the learned skills across users through inference.
  3. Monetize the model across many users without retraining each instance from scratch.
  • Anthropic's revenue growing 10x year over year while compute grows 3x year over year.

💰 Monetization

Higher AI model margins Margins could rise from around 80% for some top models to greater than 90%. 02:39

Leading labs can capture more value if their models are sufficiently better than alternatives or use compute more efficiently.

  • A model that achieves the same result with less compute can justify a larger premium.
Higher compute prices Google's rental price is 2x the spot price per hour for the GPUs discussed. 02:14

Compute providers and other companies below the labs can capture surplus by charging more for scarce, specialized compute.

  • Google paying 900 million dollars a month for 110,000 GPUs.

🔍 SEO & discoverability

Other channels

  • The author directs viewers to dwarkesh.com for other posts and future release notifications.

🧭 Frameworks

Three mechanisms for closing the revenue-compute gap00:46
  1. Increase lab margins.
  2. Increase the price of compute.
  3. Increase the percentage of compute allocated to inference.
Compute capacity growth decomposition07:51
  1. 1.4x comes from Moore's Law.
  2. 1.2x comes from building new fabs.
  3. 1.8x comes from shifting wafer allocation from smartphones and PCs to AI.

🧰 Tools & AI usage

  • Mercury — Banking platform used to review and categorize business transactions.08:57
  • Mercury Command — Built-in AI tool for proposing transaction categories and explaining its rationale.09:22
  • QuickBooks — Accounting platform synchronized with Mercury after transaction categorization.09:34

AI is used for

  • Transaction categorization — Mercury Command proposes categories and rationales for transactions so users can review, correct, and approve them in bulk.09:22

📊 Numbers mentioned

Costs

  • Google is paying 900 million dollars a month for 110,000 GPUs.
  • OpenAI spent a quarter of its compute on inference in 2024.
  • Inference may now account for closer to 50% or higher of OpenAI's compute.

Growth

  • Lab compute grows 3x year over year.
  • AI's share of leading-edge TSMC N3 wafer allocation could rise from 60% to 86% by the end of next year.

Pricing

  • Compute spot prices are more than 40% higher than the February trough.
  • Google's GPU rental price is 2x the spot price per hour.
  • A human-level software engineer running on an H100 equivalent could make that H100 worth over 250K a year.
  • The estimated H100 rental value is over 15x the current H100 spot price.
  • Renting an H100 is described at 20 dollars an hour.

Revenue

  • Anthropic's revenue grew 10x year over year for three consecutive years.
  • Anthropic ended last year with nine billion in revenue.
  • Anthropic could end this year with 100 billion to 150 billion dollars in revenue.
  • Continuing the trend would require Anthropic to reach one trillion dollars in revenue by the end of next year.

⚖️ Advantages, risks & lessons

Advantages

  • More capable models can monetize the same amount of compute more effectively.
  • Efficient models can use less compute for the same result, effectively creating more compute.
  • Training costs and learned skills can be shared across users, creating strong economies of scale.

Risks

  • Compute prices could rise enough to make it harder for other companies to compete with top labs.
  • Current popular AI applications could be priced out by higher-value uses such as AI research automation.
  • Compute supply may be unable to absorb large demand shocks because it has limited elasticity and substitutes.
  • Strong economies of scale for intelligence could increase power concentration.
  • The scarcity analysis could be wrong if market signals, ingenuity, or new supply mechanisms reduce compute constraints.

Lessons

  • When revenue grows faster than compute capacity, value must be captured through higher margins, higher input prices, or changes in resource allocation.
  • The value of compute depends on the value produced by the model using it, not only on hardware supply.
  • A model's efficiency can be economically valuable because it reduces consumption of a scarce input.
  • Compute scaling may be constrained by physical semiconductor and wafer-capacity bottlenecks.

💬 Quotes

If you have a model that can get the same result by using less compute, then you've, in some sense, created more compute, and the value of compute is gonna increase.

It captures the central connection between model efficiency, compute scarcity, and compute prices.06:11

👤 People & companies

Dylan

Podcast guest who discussed bottlenecks related to building new ASML EUV machines.

08:17
Paul Ehrlich

Figure associated with the Simon-Ehrlich bet about commodity prices and population growth.

06:54
Anthropic

AI lab whose revenue, inference margins, compute allocation, and compute rentals are discussed.

00:05
Google

Company paying 900 million dollars a month for 110,000 GPUs rented from SpaceX.

03:44
SpaceX

Company from which Google and Anthropic are renting compute.

03:44
OpenAI

AI lab referenced in discussions of compute allocation and willingness to pay for tokens.

01:24
TSMC

Chip manufacturer whose leading-edge N3 wafer allocation to AI is discussed.

08:36
ASML

Company whose EUV machines are described as a bottleneck for new-fab construction.

08:17
Mercury

Fintech company and banking platform used for transaction categorization.

08:57
QuickBooks

Accounting platform that receives synchronized transaction data from Mercury.

09:34
Choice Financial Group

Institution through which Mercury banking services are provided.

09:57
Column N.A.

Institution through which Mercury banking services are provided.

09:57

🔗 Links mentioned