YouTube Summaries

← All summaries

Nvidia's moat is memory pricing, and it's cracking

2026-08-29 Sat ⏱ 25 min t3dotgg

Theo argues Nvidia's pricing power comes less from raw compute than from artificially splitting chip performance and memory capacity into separate products you have to pay a huge premium to combine. Three things landed at once that attack exactly that: Z.AI serving GLM 5.3 Flash entirely on Huawei chips, Apple's M5 Ultra Mac Studio with 512 GB of unified memory, and OpenAI's in-house inference chip "Jalapeno". None of them kill Nvidia, but together they show the moat is narrower than the market thinks.

Everybody wants out

Nvidia's lock-in has two pillars: CUDA, which almost all AI research targets, and the chips themselves. The problem is that every large customer now sees Nvidia as an uncomfortable single point of failure. OpenAI could not function if Nvidia decided to charge ten times more; the US government's data-center buildout depends on Nvidia's cooperation. Everyone is looking for an exit, and China - largely cut off from high-end Nvidia parts by export controls - has been forced to find one first.

The memory split is the real product

A GPU is two things: the silicon that does the math, and the high-bandwidth memory that holds the weights. Nvidia prices these separately in a way that looks like deliberate segmentation:

  • RTX 5090 :: fast chip, 32 GB GDDR7, ~$2,000 MSRP (street price closer to $4,500).
  • RTX Pro 6000 :: the same speed or slower, but 96 GB of ECC GDDR7, $12,000-16,000. You are paying 6-8x for memory, not compute.
  • DGX Spark :: $4,000, 128 GB of unified LPDDR5, but only ~6,000 CUDA cores against 24,000+ on a 5090, and memory roughly a seventh the speed. It ships as an Arm Ubuntu box Theo considers unusable as an actual computer - inference appliance only.

So there is a cheap option if you need memory, a cheap option if you need speed, and a 10x markup if you need both. The obvious workaround

  • buy two cheap cards and pool the memory - does not work: consumer

hardware has no interconnect anywhere near the ~48 Gb/s the memory itself runs at. His analogy: adding fridge space by renting a building a mile from the restaurant. Mixture-of-experts models soften this somewhat, but not enough.

Chinese chips already serve frontier traffic

The anonymous "Ox Alpha" model (GLM 5.3 Flash) was served at an enormous free throughput entirely on Huawei chips, with per-token cost and hardware efficiency comparable to Nvidia GPUs. Theo had assumed the traffic must be running on Nvidia, and says so directly - the correction is the point. No CUDA anywhere in the stack.

Apple's first real AI play

The refreshed Mac Studio ends a long M3-era stall. The M5 Ultra is effectively two M5 Max dies: 36 CPU cores, 80 GPU cores, up to 512 GB of memory, and 1.2 TB/s of bandwidth. Time to first token is about 10x faster than on the M1 Ultra.

The word that matters is unified - the same memory serves as RAM and VRAM, so inference can use all of it. Against Nvidia's lineup that is roughly six times the Spark's bandwidth and only 30-40% behind a 5090, with far more capacity than anything Nvidia sells at the price. At $10k MSRP it removes essentially every reason to buy a DGX Spark, and makes the RTX Pro 6000 hard to justify for memory-bound inference. The GPU itself is almost certainly still slower than Nvidia's best, but that is now the only axis where Nvidia wins outright.

Jalapeno

Per SemiAnalysis, who were invited to benchmark it early with OpenAI engineers, OpenAI's ASIC beats Blackwell on performance per watt in nearly all scenarios - and it is a general inference chip, not a model-specific one, contrary to most press coverage. Numbers: up to 216 GB of HBM4, 700 W, 13.4 petaflops FP4. FP8 sits slightly behind GB200/300 but well ahead of H100/H200, with more memory and higher bandwidth. It hit over 700 tokens/s per user at concurrency one on DeepSeek R1, and ~1,400 tokens/s per user on Kimi K2.5 and GPT-OSS, with no speculative decoding and no prefill/decode disaggregation, and GSM8K evals on par with Nvidia silicon. Caveat: only relatively small models have been tested so far.

The design target is deliberate - OpenAI is constrained by data-center power, not budget or floor space, so tokens per megawatt is the metric that matters. Jensen Huang himself said at Computex 2026 that with a gigawatt of power, throughput per watt is revenue.

Nvidia's response

On Jim Cramer's show Huang insisted he is unbothered, citing 33 years of shipping, growing market share, and the supply chain to remain everyone's largest supplier. Theo reads it as visibly rattled. He is more charitable about the $12.9B Hugging Face acquisition: training still happens on CUDA, so funding open-source AI and pushing more companies to fine-tune their own models genuinely does expand Nvidia's customer base.

Also circling: Elon Musk's Terafab, a fab project comically larger than Giga Texas or the Pentagon and currently the only thing in his bio, despite xAI holding some of the largest Nvidia contracts (Anthropic now rents GPUs from him). Intel is apparently involved too.

Theo discloses he holds AMD stock, and pointedly does not include them in the threat list - he likes them, but they are years behind.

Takeaway

The labs are moving from being Nvidia's customers to being its competitors, and AI-assisted design is part of why OpenAI caught up faster than anyone expected. The long-run constraint, he repeats, is not chips but electricity.