NVIDIA should be concerned because OpenAI, Apple, Huawei, and other companies are developing alternatives that reduce dependence on NVIDIA, although NVIDIA still has major advantages in CUDA, supply chain, technology, and market share.
🔒 17 more in the full analysis
🔒 2 more in the full analysis
Searchable transcript of NVIDIA Should Be Scared — Theo - t3․gg (25:28). Search for a phrase, then click its timestamp to jump straight to that moment in the video.
Captions sourced from the original video on YouTube, published by Theo - t3․gg. The video, its captions and all related intellectual property remain the property of their respective owners; AINotes claims no ownership. Provided for research, accessibility and search — see the Transcript Notice and Copyright Policy.
00:00 Nvidia's a weird company. They originally started making chips for gamers to have fancy 3D graphics in their computer games. And now they are the most valuable company in the world, powering the majority of intelligence that we get from all the fancy AI tools we use every single day. There's a lot of pieces that make Nvidia's borderline monopoly on compute for intelligence just impossible to crack.
00:21 things like CUDA, which is the language and system of choice for the vast majority of AI research, to the crazy deals that they're cutting with everybody from the government itself in the United States to businesses like, of course, OpenAI, Anthropic, XAI, and more. And it seems like everyone is realizing that Nvidia is a not great core dependency to have on your business's success.
00:42 OpenAI could not function if Nvidia decided they wanted to start charging them 10 times more. The US government would be pissed if their whole plan to build this crazy set of data centers fails because Nvidia changes their mind. A whole of these businesses are struggling to work with the terms Nvidia is giving them today. And with the promise that those terms will get worse tomorrow, they're all looking for ways out.
01:03 But there's one particular place that is more motivated than anywhere else to solve this. China. Nvidia chips have largely been banned from being sent to China by the US government. Some have been cleared for export, but it's a very, very small set, and it's not the giant powerful GPUs that we are using for training in the US. All of this has pretty much forced the Chinese labs like ZAI to rely heavily on chips that are being made by Chinese companies like Huawei.
01:30 But it's not even just the Chinese labs anymore. Even OpenAI has started developing their own chips to get faster, cheaper inference and training available to them internally. All of this has Nvidia acting pretty scared and very very weird. From crashing out on Jim Kramer's show to buying Hugging Face. Little bit absurd, but as you guys can probably guess, I have a lot of thoughts on this.
01:55 I'm a big nerd about chips. I'm a big nerd about Nvidia in particular, and never been the biggest fan of them. And the opportunity to point out how their monopoly is about to be crushed and crumbled is one I have to take. But unlike my friends who invested in Nvidia early, I don't have billions of dollars to spare. So, I'm going to have to take a quick break for today's sponsor.
02:12 I got a hot take for you guys. The knowledge in existing LLMs is nowhere near enough for us to do real work and get our jobs done. Thankfully, the labs have noticed this, too. And that's why the majority of the context that your agents work with isn't stuff that they already knew. It's stuff that they got from the internet. The vast majority of the context in our real context windows is things that the agents fetched from online.
02:34 What I'm trying to say is that your agents need access to the internet. And if they want the best possible way to do it, you should probably use today's sponsor, Browserbase. They provide all of the pieces your agents need to be smarter and have modern knowledge from their search API, which is literally a single curl request that allows your agents to look things up and find real context from the web to their fetch API, which lets agents give a URL to browser base and get back markdown they can actually parse and
02:57 understand from pretty much any website across the entire web. And don't forget about the browser as a service, which allows your agents to actually navigate the web like a human would. With a real Chromium browser, allowing them to fill out forms, sign into pages, and do real actions on users behalfs. Over 85% of the APIs on the web can't be accessed by simply curling them.
03:16 If you're okay with only having that 15%, stick to curl. But if you want to unblock your agents and give them the whole web, do it today at soy.link/browserbase. Before we can talk about how Nvidia loses, it's important to understand how they won. There are a couple core pieces, but I want to just fixate on the two that I'm most interested in right now.
03:35 Those pieces are, of course, the chips that are being used by all of these companies that Nvidia makes. But the other piece is CUDA. We'll come back to this one in a bit because this is where things will get complicated, but for now, I want to focus on the chips side. GPUs are uniquely tailored for these types of AI tasks because they have millions of small cores instead of a handful of big ones.
03:57 In order to traverse these gigantic piles of weights and parameters that a model is, you have to trail all of the data inside of it to make a response. You get a lot of value from having a chip that has lots of different processes on it that can handle that type of large amounts of data and crazy matrix transform [ __ ] Lots of small brains are a lot more powerful than one giant one in these cases.
04:20 You also need to have enough memory to hold the weights that are being traversed, which is another real fun problem that Nvidia is consistently causing for themselves. The two pieces of a given I'll change this to be GPUs. The two pieces that matter for a GPU are the actual like process itself. I'll say the chip and the high bandwidth memory are probably the easiest way to do this.
04:43 The GPU has these two parts. the actual silicon that has all of the things on it that can process all of this data and then the high bandwidth memory which actually holds said data. Nvidia has had a lot of fun with this split and arguably using it to fix prices in their favor. I won't say it's cheap, but getting a powerful Nvidia chip is not the hardest or most expensive thing in the world.
05:06 One of the highestend Nvidia GPUs available in terms of its actual like GPU performance and throughput is the RTX5090. The retail price for the RTX5090 was originally $2,000. They now consistently go for around $4500 because the demand is so absurd. There's also a separate card they put out called the RTX Pro 6000 that ranges between 12,000 and $16,000.
05:32 You would assume this chip must be way, way faster than the 5090 if it's going to be six to 10 times more expensive. But guess what? It is the same speed or slower. This might sound crazy. You might be confused. How is the chip that costs 6 to eight times more slower? Why would I ever get that instead of the 5090? Well, it turns out the chip isn't the only thing that matters.
05:55 The 96 gigs of GDDR7 are what mattered here. The 5090 only gets 32 gigs of GDR7 and it's not error correcting. Technically, the RTX Pro has more CUDA cores, but I believe the process on the 5090 is slightly like newer and better. It should be within spitting distance for the raw performance and throughput there, but the RAM difference is why they're able to charge so much more.
06:18 And just to be explicitly clear here, they were doing this way before RAM got more expensive. They do this because they know the people who need a fast chip and a lot of RAM are much more willing to spend lots of money than a video game player is. And we're already seeing my favorite question in chat, which is why not just connect them. I'm going to use my favorite analogy I always use for chip related stuff here, a kitchen.
06:46 Imagine you have a kitchen that serves a bunch of customers at your restaurant. You have three really talented chefs in there, but your freezer is full and your refrigerator is full. You're running out of space. So, you decide you need another fridge. You're out of space in your restaurant, though, but you need that other fridge pretty bad. So, you buy a building a mile away and you put the fridge there.
07:10 Maybe you put 10 fridges in five giant deep freezers there. You massively increase the amount of storage you have. But now every time the chef needs something from those fridges or freezers, they have to run a mile to the other place, grab it, and then run back. You can't just plug the chips in together. If you have a model that is 40 gigs, hell, if you have a model that's 33 gigs and it doesn't fit on your 5090, you now have to split the data across the two GPUs, which means ideally you have some way of predicting
07:42 which data is needed when and how so that GPU one only has the data it needs and two only needs the data that it needs. Good luck. Have fun. Not [ __ ] happening. There are techniques around things like mixture of experts models that allow you to more easily some amount assign the work across stuff, but there is no way on consumer hardware to get reasonable bandwidth between 25090s to allow them to share memory.
08:06 By the way, that memory that they're using for this is as much as 48 GBs per second. That means you need something even faster in order to transfer the data between the GPUs. Once again, not happening. Don't worry, though. Nvidia has an answer for you. If you need more memory, they're more than willing to sell you something. And no, I'm not referring to the RTX Pro.
08:29 They know that there are people who can't put out 18 grand plus on a GPU that they're just going to use to run shitty models locally. And that's why they made the DGX Spark. This small little box costs four grand. And I hope you don't plan to use it as a computer. That is what it is to be clear. It's a computer. It shows up running a botched install of Ubuntu that they filled with [ __ ] and it runs a 20 core ARM chip.
08:56 And having spent far far far too much time fighting ARM Linux in my life, I promise you, you're not going to use this computer as a computer. Linux on ARM is hell. Don't bother for this for anything other than inference. So, this ARM chip in this random $4,000 mini PC that you can't use as a computer does have a benefit. The 128 gigs of unified memory.
09:23 That memory is LPDDR5 memory, though, which says a different number here. There's a lot of layers to how those numbers are measured. LPDDR5 is meaningfully slower than GDDR7. So, the RAM is slower, but you have way more of it. But most importantly, you have CUDA cores. cores that can be used with CUDA backed stuff that are actually capable of running real workflows, but you get 6,000 of them instead of the 24,000 plus you get on hardware like a 5090 or an RTX Pro.
09:56 So, your options here are the world's worst computer environment possibly that you can buy today for real amounts of money, but you get a real amount of RAM you can use for models at the cost of an absolute [ __ ] garbage chip that runs terribly slow. Or you can get a really powerful capable chip like an RTX5090 and now you're entirely gimped on RAM.
10:18 And if you want RAM and a high-end chip, you're paying the $16,000 plus dollars for that RTX Pro sadly. Do you see what they've done here? They have the cheap option if you need more RAM. They have the cheap option if you want a fast chip, but you're paying up to 10 times plus more if you need both. They did do one really nice thing with the Sparks.
10:37 They gave them a 200 Gbit per second nick so that you can connect them over SFP so you can have the different sparks transfer data between each other relatively fast. It's hell to set up, but you can at least do it. So, we've addressed all the fun things here for where Nvidia is at and how they can charge these absurd prices. What we haven't yet addressed is why I'm filming this today.
10:59 There's a couple things that inspire me to take this video on. The first is an anonymous model that dropped last week called Ox Alpha. I already have a video about this is probably already out on the channel by the way. You should check that out. And while you're on the way to it, you should hit that subscribe button because it cost you nothing and these videos take a lot of work.
11:15 So OX Alpha was an anonymous model that came out on Open Router and Open Code, both of which had an absurd amount of free throughput on them. I believe it was 100 trillion tokens a day for free. And the model was great. A bunch of people were using it. I was using it a bunch. I have a video about how much I love it. Great model. Turned out to be GLM53 Flash.
11:34 And the reason they could serve it so aggressively is both because it's a relatively small model, but also because they served it entirely on Chinese chips. Huawei made the chips that they served all this traffic on, and it's working great. It's genuinely impressive that without any Nvidia in their stack, they were able to do this absurd level of traffic.
11:51 I also got called out here because people were saying only Frontier Labs have this much compute. When I said that, I assumed that they were still serving Nvidia because we basically expected that and also this wasn't a small model. To be fair, when I tweeted that, my secret personal belief was that Xiai was doing this on Chinese chips somehow. Then it turned out to be GLM shipping again in like a twoe window, also on those same chips.
12:17 All of the traffic was served on Chinese chips, attaining hardware efficiency and per token cost comparable to Nvidia GPUs. The CUDA mode is being tested once again after Jalapeno's announcement yesterday. Jalapeno is OpenAI's chip that they've been working on in order to get out of the hell of relying on Nvidia so much. They are partnering with Cerebrris, who is a company that makes faster inference chips and hosts models that way in order to do everything they can to massively increase speeds and reduce costs.
12:41 But that is not the only thing that was just announced to scare Nvidia. Apple out of nowhere announced upgrades to the Mac Mini, but more importantly the Mac Studio, which has not seen an upgrade since the M3 era. And the M4 did kind of come out, but they never made an M4 Ultra. They only made an M4 Max, and it was gimped on how much RAM it could have.
13:02 So, you were stuck with like a 3 and a half plus year old chip if you wanted a lot of RAM on a Mac Studio. Obnoxious, dumb, solved. The Mac Studio now has M5 Max and Ultra. The first section here is for the max, which should do up to 128 gigs of RAM and 614 GB per second of memory bandwidth. But the Ultra is basically just two of these chips stapled to each other, which means it can do 36 cores of CPU, 80 cores of GPU, up to 512 gigs of memory, 1.2 terabytes a second of bandwidth, and a 32 core neural engine.
13:36 As a video nerd, I love the fact that the new Ultra chip can do 33 streams of 8K ProRes 42:2 at 30 fps playback. Like, that's insane. But that's not what we're here for. Let's be real. We are here because the time to first token on a Mac Studio with M5 Ultra is 10 times faster than it was on a Mac Studio with M1 Ultra. That is a massive increase in performance.
14:01 And this largely comes down to the improvements in the memory. But the important thing to note here with the memory isn't even just the amount. It's this word here unified because that means it works as normal RAM but also as VRAM which means the GPUs can use it for inference. When you look at the Mac Studio compared to Nvidia's options, you see just how compelling it gets.
14:21 You can get a 5090 which has crazy GPU performance which isn't indicated here how fast the compute is on it. You only get 32 gigs of RAM though at the benefit of the way Nvidia implements it. Effective 1,800 gigabyte per second memory bandwidth. The RTX Pro, similar bandwidth, 3x the RAM, much more than 3x the price. The DGX Spark, way more RAM up to 128 gigs, but it's unified LPDDR5.
14:52 That is a seventh or so the speed. But suddenly we get a thing with no compromises. No compromise on the amount of RAM, no compromise on the memory bandwidth, and depending on how it performs when it comes out, not as big of a compromise on the compute itself. So for 10 grand MSRP and the current price, still that because it is coming out soon, you get way more RAM than any other option that's offered by Nvidia.
15:16 A chip that is meaningfully better than the DGX Spark, comically so even not necessarily better than the GPUs in the 5090 and the 6000 Black. Well, we don't know yet until it's actually out, but I'm guessing almost certainly not. But most importantly, you get 256 gigs of their unified memory with bandwidth a hell of a lot closer to what Nvidia sees on their GPUs than to what that you would get on a Spark.
15:39 Almost six times faster than the Spark and only like 30 to 40% slower than the 5090. That is insane. This kills almost all of the reasons you could ever justify buying the DGX Spark, which is a whole category of Nvidia devices killed. And to all the people who upgraded from a Spark to a RTX Pro 6000 or skipped the Spark because they wanted speeds that weren't trash, they can now get way, way better experiences for meaningfully cheaper by just getting the M5 Ultra instead.
16:13 That's crazy. This is Apple's first real play in the AI space. Them coming in and saying, "Sorry guys, you're [ __ ] around too much. We're going to put an end to that." And yes, the DGX Sparks effective memory bandwidth is actually this pathetic. It's a [ __ ] chip. It's a [ __ ] system. The DGX Spark is my least favorite computer in this apartment.
16:33 But what about Jalapeno? Well, according to semi analysis, it is coming out to be better than Blackwell. Their self-designed ASIC is comparing with Reuben, which is Jalapeno's TCO through per megawatt and spicy deetss. Cool. OpenAI actually invited semi- analysis to come take a look early. In general, first generation trips are not competitive, but OpenAI bucks that trend by being industry leading in beating every NVIDIA, AMD, and Google chip we have been able to test on multiple top open source models.
16:59 OpenAI does this with extreme hardware, software, code design. Traditionally, these bespoke chips tend to super fixate and specialize on specific things. They have not actually done that. They built a really good general chip that delivers high performance in all scenarios. Apparently, their chip goes as high as 216 GB of HBM. It only requires 700 watts.
17:20 It's doing 13.4 pedlops for FP4. And yeah, those are insane numbers. Apparently, it doesn't have FP16 numbers. But for FP8, it is slightly behind what you're seeing on GB2000 and 3000s from Nvidia, but it's meaningfully ahead of the H100 and 200 already, which is pretty crazy. And also more RAM and way higher bandwidth. Even just this chart should be enough to give Nvidia a heart attack.
17:45 Still not quite Reuben levels, but it's trading blows at a way lower wattage. A lot of the media coverage of this chip has followed a few throwaway comments from OpenAI that claim the chip will be optimized for their models in a way that other chips aren't. This is wrong. Jalapeno is a generalized inference chip capable of running all sorts of models in all sorts of workloads, including our benchmark inference X where we ran the benchmark with OpenAI engineers in the lab.
18:08 As a joke, OpenAI even showed us it running Doom, which was ported to their chip with just codeex prompts. Of course, they have it running Doom. The following is the headline performance per watt result. So let's see what the numbers looked like. This is looking pretty insane. Token per second per megawatt performance here. They are crushing everything else in efficiency in terms of electricity.
18:28 Jalapeno is beating Blackwell on performance per watt across almost all scenarios without being tuned for any specific point in the curve. It excels not only at low latency scenarios but also in high throughput scenarios. A more applesto-apples comparison is against single token prediction results. It knocks every competitor out of the water. At low concurrency scenarios, Jalapeno demonstrates remarkable interactivity, hitting over 700 TPS per user at concurrency 1 on the DeepSeek R1 model.
18:54 This is all being achieved with single token prediction, no speculative decoding, and no pre-filled decode disagregation. They got Kimmy K25 and GPT OSS running at 1,400 tokens per second per user. And they confirmed that the GSM AK valves attained results on par with the video chip. So they're not nerfing the models when they run them. Since they're using HBM4, which is the new generation of high bandwidth memory, they get a pretty substantial win over things on older memory.
19:18 Reuben is the new Nvidia line that will use HBM4, but those chips are not really actually out yet. Blackwell is the line that most things are buying and using. There are people who have put in huge orders for Blackwell chips that will finally show up in like two to three years. Is a while before Rubin's going to matter. There's also the call out that this is not large model performance.
19:37 All the models they've tested so far are relatively easy to run in terms of size. So we don't know how these will perform when you give them huge models like Kimmy K3 or DCP4 Pro. OpenAI is designing for performance per watt. The reason is simple. OpenAI is currently limited by data center power, not by budget or floor space, and thus tokens per megawatt is paramount.
19:56 At Computex 2026, Jensen said the performance per watt, reliability, and long lifetimes are the core features of future GPUs. to quote, "If you have one gigawatt of power, then throughput per watt is revenue." Yep, I've been talking about this for a while. Electricity is going to be a big deal. So, OpenAI focusing on that side is a huge deal. Don't worry, though.
20:18 Jensen's definitely not scared. Nvidia has a plan. They're going to be totally fine. Not like he's going to say something stupid like they're just going to cancel the development of the thing that destroys our business model. I haven't actually seen this clip, so we get to watch it together. >> And I said, "You know what? I the whole time I've been working on on something a chip that I think could really hurt you.
20:37 It's called jalapeno. I don't know. Let's call it I don't know tomatillaa. And you said to me, well that's fine. Would you really say that's fine? Because I personally would be hurt if open AI did to what they do to Nvidia if they did to me. >> You know, I'm okay with it, Jim. There's so many XPUs that are being announced and as we know, it's not easy doing what we do.
20:58 We've been doing this for 33 years and so lots of projects get started, a lots of projects get gets cancelled. Um we're we're here we're here to support our partners and and um we're going to build the world's best technology. I have every confidence in that. Uh we're going to be the most productive infrastructure that they have. I have every confidence in that.
21:19 Um we have the supply chain and the technology scale to be their largest supplier. I have every confidence in that. And so, you know, I I don't I don't have to take anything personally because I've got so much confidence in what we we're able to deliver. And look at look at all of the XPU announcements and all the startups that have been announced. And yet today, Nvidia is increasing our market share of the AI market.
21:42 We're our growth is accelerating. Our technology leadership is extending. And so, I'm very comfortable with all the competition. >> Okay. And >> yeah, definitely not scared at all. Are you bud? Well, on the bright side, if they own Hugging Face, they can make sure they suppress access to all of the versions of the models that run well on things that aren't CUDA.
22:03 I will say that I think the Hugging Face bid is an actual good faith play to try and bolster and fund the development of open-source AI and encouraging more and more people to train because let's be real, training is still happening on CUDA. The more they encourage businesses to try and train and fine-tune and customize things themselves, the more customers they have for their chips, the longer term their absurd saturation can go for.
22:27 And Nvidia is still making a ton of money. They can justify doing [ __ ] like this. Good for them. If you want to spend less than 12.9 billion to do something that positions your business better, you have access to my DMs. Jensen, somebody in chat mentioned, "You know what's funny, Theo? I bet SpaceX AI is also doing something similar now." which reminded me somehow I entirely forgot about Terraab, the most epic chip building effort ever, which is by the way the only thing in Elon Musk's bio right now, despite the fact
22:54 that him and Jensen are buddy buddy and Elon has some of the biggest and most lucrative contracts with Nvidia. He has more GPUs than Anthropic does. Anthropic is renting GPUs from Elon now because he was so quick on this and he is still concerned about Nvidia's monopoly and just trying to get out of it. Terrafab will close the gap between today's chip production and the future's demand.
23:14 A future among the stars. And we do this by building a gigantic chip fab. Terapab will be comically bigger than even giant things like the Giga Texas fab for Tesla, the US Pentagon, Mall of America, and more. It's 25 times the size of Apple Park. It's 20 times the size of the Pentagon. Kind of crazy. So yeah, I guess you could say Elon and SpaceX AI and Tesla or whoever whatever business he has doing this are considering competing with Nvidia.
23:49 It's almost like literally everyone is. I should include a call out here. I am an AMD investor. I am not a special early investor. I'm just buying their stocks, but I have not talked about AMD at any point here because I'll be real. I love them to death. They're very behind. It'll be a while before they can catch up to what these other things are that I'm talking about here.
24:09 I hope that changes, but for now, AMD is a potential big winner, but they have some time before they get there. I think that's all I have to say on the chaos that is the current state of chips. Apparently, Intel is involved in Terra Fab 2 fun. That'll be an interesting project. I have no idea where any of this will go. All I know is that Nvidia is scared and they have good reason to be.
24:29 The AI world constantly is changing and these tools and technologies make it easier than ever to catch up. We finally now have models that are useful enough to help these manufacturers in their process. I think this is a big part of why Open AI is catching up as quickly as they are. Their use of AI in their catch-up process has enabled them to do it more effectively than anyone would have anticipated, including Nvidia.
24:48 And now we're quickly approaching a future where the labs are able to compete with Nvidia directly instead of relying on them with these trillion dollar contracts to get all of the chips they need. As I've said many times now, the future is going to be fought not on chips, but on electricity. And if we don't have ways to get the energy we need, then none of this ends up mattering in the end.
25:08 But at the very least, for now, it is super interesting. And Nvidia's weird position in the market might not last as long as they probably think. Am I crazy for saying all of this, or am I kind of on to something? Let me know how you guys feel about the future for Nvidia and this whole space in general in the comments. And until next time, peace nerds.