holocron
← Business
Business · acquired2023-09-06

NVIDIA Part III: The Dawn of the AI Era (2022-2023)

In one sentence: Part three of Acquired's NVIDIA trilogy (Season 13, Episode 3, recorded and released 2023-09-06 — every number and judgment on this page is as of that moment and carries nothing that came after). Only 18 months after Parts I & II, the show had to reopen the story: Ben ran a transcript search and the April 2022 episodes never once said the word "generative" — "That is how fast things have changed." The episode's thesis: the 2021 $1T TAM slide stitched together from autonomous vehicles, robotics, and the Omniverse basically didn't come to pass; what came to pass was a market that wasn't on the slide at all — LLM training. David's "unintentional and uninformed" throwaway question at the end of Part II — what if the digital world grows a new foundational layer that Nvidia can power? — came true. Ben: "In November of 2022, AI definitely had its Netscape moment. And time will tell, but it may have even been its iPhone moment." David: "That is definitely what Jensen believes."

The Company on One Page (this episode's timeline)

YearEvent
1960sConvolutional neural networks already existed, but training them was computationally out of reach — "Nobody thought it would be practical to actually train and use these things, at least not anytime soon or in our lifetimes."
2006CUDA launches (Jensen, Ian Buck, and others), a bet on scientific computing; the same year, the decision that every GPU shipped would be fully CUDA-capable — an expensive call grumbled about internally for years, genius in hindsight
2012AlexNet (Jensen: AI's "big bang moment"): the Toronto trio (Krizhevsky / Hinton / Sutskever) used two ~$1,000 consumer GTX 580s plus CUDA to cut ImageNet's mislabeling rate from 25% to 15%; acquired by Google within 6 months into the freshly formed Google Brain
2013–14Google buys DeepMind, Facebook lands Yann LeCun; the top-researcher duopoly forms; Google Brain redoes YouTube recommendations, Instagram becomes a $100B–$500B asset for Meta on AI feed recommendations — the first AI monetization wave is recommendation-feed advertising, and the giants start buying Nvidia GPUs en masse
2015Musk and Altman convene a fateful dinner at the Rosewood Hotel on Sand Hill Road to poach researchers, and pry loose exactly one — Ilya; OpenAI is founded as a nonprofit; Karpathy publishes his language-model vision blog post
2017Google Brain publishes the Transformer paper, "Attention is All You Need"; in the window that follows, only Google grasps its weight
Fall 2018Musk leaves OpenAI, which then pivots fully to Transformers — David: "a history-turning-on-a-knife-point moment," and "probably super bad decision by Elon"
20193/11 OpenAI converts to capped-profit; within 6 months Microsoft invests $1B as exclusive cloud provider; in March Nvidia announces the $7B cash acquisition of Mellanox (blowing Intel out of the water); in August it ships Megatron: 8.3B parameters, 512 GPUs, 9 days of training, ~$500,000
2020June: GPT-3; September: Microsoft licenses exclusive commercial use of the model; Mellanox closes — David: "one of the best acquisitions of all time, and nobody had any idea"
2021GitHub Copilot ships; Microsoft invests another $2B; Nvidia puts up the $1T TAM slide ($100T × 1%)
2022-04Acquired publishes Parts I & II; Nvidia is the 8th most valuable company in the world, ~$660B
2022Rates spike; Ethereum moves to proof of stake, ending mining demand; a huge write-down on pre-ordered TSMC capacity; market cap crashes below $300B; September brings the Biden administration's China export controls (spawning the nerfed A800/H800 SKUs) + the Grace CPU announcement + the architecture split (data-center Hopper/H100 vs consumer Lovelace)
2022-11-30ChatGPT — Jensen: "the AI heard around the world"; the fastest product in history to 100 million users
2023-01Microsoft's third round, $10B, and GPT integrated across all products
Spring 2023At GTC, Jensen's fireside chat with Ilya (source of the "detective novel" argument; David strongly recommends watching the whole thing)
2023-05FY24 Q1: revenue $7.2B (+19% QoQ), plus the bombshell $11B Q2 guidance; the stock pops +25% after-hours, jumping from ~$800B into the trillion-dollar club (GPT-4 is narrated as shipping "in May of this year," actually 2023-03, recorded as said)
2023-08FY24 Q2, reported the week before recording: revenue $13.5B (+88% QoQ, more than doubling YoY); the data-center segment alone did $10.3B in the quarter (+141% QoQ, +171% YoY; a segment that barely existed 5 years ago); gross margin 70%, next-quarter guide 72% (pre-CUDA it was 24%) — David: "one of, if not the most, incredible earnings release by any scaled public company ever"

Founder Profile

Jensen Huang (age 60)

The judge. He calls AlexNet AI's big bang, ChatGPT AI's iPhone moment (a framing he firmly believes), and generative AI a "new computing era." His core sermon is "the data center is the computer" — in Ben's paraphrase of his tone: "hey dummies… Listen to me when I tell you the whole data center needs to be one computer." The new trillion-dollar narrative on the Q2 call: $1T of hard assets are already sitting in data centers worldwide, with $250B of new CapEx a year. He asserts every application will have a GPT front end — which Ben reads not as "replacing button clicks" but as everyone becoming a programmer, where the programming language is English. His strategy self-description (to Ben Thompson): "We build our systems full stack, but we go to market in a disaggregated way." His operating philosophy: "You build a great company by doing things that other people can't do. You don't build a company by fighting other people to do things that everyone can do."

Management and routine. Rumored 40 direct reports; his office is basically an empty conference room because he's bouncing around constantly. He doesn't talk to senior execs about their career ambitions — "they're the best in the world at what they do. This is their life's work." As he put it: "I start my day at 5:00 AM seven days a week. You do too." Ben judges he "has another 30 years in him" and has built a company with no succession pipeline — "the company is an extension of Jensen's thoughts, will, drive, and belief about the future." Asked what he does to relax, he answered 1000% seriously:

"I relax all the time. I enjoy relaxing at work because work is relaxing for me. Solving problems is relaxing for me. Achieving something is relaxing for me."

He isn't buying sports franchises or mega yachts — "or if he is, he isn't talking about them, and he is working from them" (Ben's kicker: he's also not buying social media platforms and newspapers). The keynote is "the Jensen show," one man start to finish (contrast Apple, where Tim Cook does the welcome and hands off to a parade of executives); every time, he makes a big deal of the H100's weight — "oh, I can't lift it" (70 pounds) — and his signature line "the more you buy, the more you save," paired with the energy argument: these machines burn enormous power, but the alternative burns even more.

Ben adds a second reading of the "iPhone moment": the iPhone is fundamentally a hardware company differentiated by software that then expanded into services — which makes Nvidia's Apple analogy uncannily apt (a vertically integrated hardware-software stack; developers are most motivated to target the highest-volume platform), and its customers are the least price-sensitive B2B buyers of all. The one difference: "Apple has always had a market cap that lagged its proven value to users. Whereas Nvidia right now is exactly over their skis."

The AI Cast (where this episode concentrates its cast of characters)

  • Geoff Hinton: the legendary computer-science professor and AlexNet faculty advisor; great-great-grandson of George and Mary Boole. Ben: "This guy was born to be a computer science researcher."
  • Ilya Sutskever: the third member of AlexNet; followed the team into Google Brain and worked on redoing the YouTube algorithm; at the 2015 Rosewood dinner he was the only one intrigued, and left to co-found OpenAI as chief scientist — per Wired's Cade Metz, in his own words: "I felt there were risks involved, but I also felt it would be a very interesting thing to try." He supplies the detective-novel argument at Spring 2023 GTC (see Playbook #5).
  • Andrej Karpathy: author of the 2015 blog post The Unreasonable Effectiveness of Neural Networks (Ben's rendering of the title; the post is actually about recurrent neural networks, recorded as said); his arc runs OpenAI → head of Tesla AI → back at OpenAI. In a 1-minute-45-second Nvidia-channel video (2015/16), two years before the Transformer, he already sketched it: "The idea that you can take a large amount of data, and you feed it into the network, and it figures out the pattern in how words follow each other in sentences… you can train basically a chat bot… Eventually we'll use that to talk to computers just like we talk to each other." David: "Wow. This is 2015?"
  • Elon Musk and Sam Altman: the two who, in 2015, saw the duopoly as a serious problem (Altman was then president of Y Combinator). Ben on OpenAI's pedigree: "It is very different from this organic, scrappy way that the Nvidias of the world got started. This is powers on high and existing money saying, no, we need to will something into existence." Musk's 2018 exit was a major catalyst for OpenAI's pivot; Sam drove the for-profit conversion and the Microsoft deals — David's two-sided read: a skeptic would say "okay Sam, you took your nonprofit and you converted it into an entity worth $30 billion today," but knowing the history, "this was the only path they had."
  • The Google Brain founders (Greg Corrado / Jeff Dean / Andrew Ng) and Yann LeCun: the former turned AlexNet into enormous Google profits; the latter, hired by Facebook, anchored the other pole of the duopoly.
  • Ian Buck: co-launched CUDA with Jensen in 2006, now runs the data-center business, and was interviewed for this episode.
  • The hosts as their own material: David's accidental prophecy at the end of Part II — "What if maybe, just maybe, the Internet, software, and the digital world are going to keep growing, and there will be a new foundational layer that Nvidia can power?" — "I can't believe that I said this because it was unintentional and uninformed."

The Playbook

Each entry: origin story → insight → effect.

1. Accelerated computing is the Archimedes lever on Moore's Law

  • Story: AlexNet did work that used to require a supercomputer lab using two ~$1,000 GTX 580s you could buy at Best Buy, with the algorithm written in CUDA.
  • Insight: a CPU executes one instruction at a time; a GPU executes hundreds or thousands. As long as a problem is parallelizable, you can lever up Moore's Law by hundreds, thousands, and now tens of thousands of times — David: "a giant Archimedes lever."
  • Effect: a parallel processor built for graphics turns out to fit every linear-algebra workload; AI and crypto are new frontiers for the same technology — spillover Nvidia never originally foresaw.

2. The von Neumann bottleneck is the key to the GPU revolution — and the constraint has now migrated to memory

  • Story: Ben does his "best professor impression" live: computing 2+3=5 takes four instructions — load, load, add, store — and three of four clock cycles go to reading and writing memory; even a 1 GHz CPU can run that tiny program only 250 million times a second.
  • Insight: adding memory doesn't help, and raising the clock speed only helps linearly — worse, the bottleneck deteriorates over time. The unlock is to abandon von Neumann: parallel execution plus massive core counts. After parallelizing, the new constraint is memory: models run hundreds of GB and must all be resident, yet the H100 has only ~80 GB of on-chip RAM, and EUV reticle limits mean chips physically can't be made bigger.
  • Effect: you must network multiple chips, servers, and racks into "one computer" — "huge amounts of memory, very close to the processors, all running in parallel, with the fastest possible data transfer" — one sentence that explains everything about Nvidia's data-center strategy.

3. The Transformer's real breakthrough = trading parallelism for context

  • Story: RNNs/LSTMs are sequential with a very short attention span — "like a human with a very, very short attention span" (translating "The United States" to Spanish fails on the very first word: Estados Unidos); attention lets the model look back over the entire input and weigh it when picking each output word, at a cost of O(N²) — double the input, quadruple the cost.
  • Insight: those comparisons can all be done in parallel; with enough GPU cores, a thousand-word input takes about as long as a ten-word one — the first time sequence models could be trained in parallel, at a scale that simply couldn't be trained before.
  • Effect: it unlocks LLMs; the quadratic cost curve means every step up in model size explodes GPU demand.

4. Knowledge is embedded in the data: unsupervised pre-training and the emergence of scaling

  • Story: GPT-1 (120M parameters) proved you can learn language structure from unlabeled text — "like how a child consumes the world where only occasionally does their parents say, no, no, no… that's actually the color red." Then GPT-2 (1.5B), GPT-3 (175B), and GPT-4 (unannounced, rumored ~1.7T). Ben: "It's scaling like Nvidia's market cap."
  • Insight: models below 10 billion — maybe even 100 billion — parameters mostly just hallucinate; without changing the structure, merely adding data and parameters makes the output "magically" better — "No researchers expected them to reason about the world as well as they do, but it just happened as they were exploring larger and larger models."
  • Effect: huge training data ⇒ huge amounts of Nvidia GPUs; unsupervised pre-training followed by supervised fine-tuning becomes the standard paradigm.

5. Prediction is understanding (Ilya's epistemology)

  • Story: at Spring 2023 GTC, Jensen deliberately floats a straw man — LLMs "are just probabilistically predicting the next word… They don't actually have knowledge." Ilya's response: consider a detective novel where, at the end, the detective gathers everyone in a room and says, "I am now going to tell you all the name of the person who committed the crime. And that person's name is blank."
  • Insight: the more accurately the model predicts that word, ipso facto, the deeper its understanding not just of the novel but of all general human knowledge and intelligence; today's GPT-3, GPT-4, Llama, and Bard can all guess the criminal.
  • Effect: Ben flags it — "Put a pin in that, understanding versus predicting, the hot topic du jour" — an epistemological defense of the scaling path.

6. Compute cost is the new moat: the only logic behind OpenAI going for-profit

  • Story: after the Transformer, the roadmap was visible to everyone (OpenAI didn't have its head in the sand), but training costs were untenable for a nonprofit — and for nearly everyone except Google. On 2019-03-11 OpenAI converted to capped-profit — its own words: "no pre-existing legal structure that we know of strikes the right balance."
  • Insight: for now, the lead is decided by CapEx capacity — whoever can afford the compute earns a seat at the table.
  • Effect: three rounds of Microsoft money bought exclusive cloud plus an exclusive commercial license; and OpenAI, as far as anyone knows, runs entirely on Nvidia infrastructure, just purchased through Microsoft — whichever way the structure shifts, Nvidia is the revenue beneficiary.

7. The Mellanox acquisition was research-driven foresight

  • Story: in March 2019 Nvidia paid $7B cash for a "quirky little" company nobody understood; five months later it shipped Megatron (8.3B parameters, 512 GPUs, 9 days) — internal research had already convinced them models would run across servers and racks, making inter-machine bandwidth the crux. Ben: "I didn't discover this until two hours before we started recording."
  • Insight: InfiniBand is far faster than ethernet, but the industry consensus that "ethernet is the lowest common denominator" made vendors exit, leaving Mellanox as the only InfiniBand-spec provider left; the point of 3,200 gigabits/second between racks is to treat memory in a box three feet away as if it were near-on-chip memory.
  • Effect: David: "one of the best acquisitions of all time" — others couldn't see it precisely because they hadn't done the research.

8. The data center is the computer: build full stack, go to market disaggregated

  • Story: in September 2022 Nvidia announced the Grace CPU (an Nvidia CPU had been considered heretical). Ben: this is "the end game of a ballet that has been in motion for 30 years" — the graphics card that once sat subservient to a PCIe slot on Intel's motherboard now becomes "our GPUs, our CPU, our NVLink between them, our InfiniBand to network it to other boxes. Welcome to the show."
  • Insight: chip-to-chip comparison isn't the arena; the battlefield is how multiple GPUs and racks work together as one system (hardware + networking + software) — Nvidia changed the vector of competition. Three pillars: the Mellanox networking stack, the Grace CPU (off-the-shelf ARM with Nvidia's own tuning, not an Apple-M3-style build-from-scratch hero product), and the purpose-built Hopper architecture; the dream is a customer owning and operating a DGX SuperPOD, the reality is components that can be sold piecemeal or in the cloud.
  • Effect: there's nothing else on the market like it; they're effectively selling mainframes — "it's not that different from IBM way back when they're trying to sell you a $100 million wall that goes in your data center."

9. The architecture split is a CoWoS capacity-monopoly strategy

  • Story: the data center goes Hopper — using TSMC's bleeding-edge CoWoS (chip-on-wafer-on-substrate) 2.5D packaging to stack HBM; the gaming line is split off to Lovelace, which doesn't use CoWoS.
  • Insight: CoWoS is only 10%–15% of TSMC's total capacity, and adding more means building new fabs over years; the split lets Nvidia monopolize a large chunk of CoWoS capacity for H100s — the truth behind "the chip shortage" is a packaging-capacity bottleneck.
  • Effect: the H100 carries far more memory than any other GPU on the market; even if a competitor's chip design catches up, it can't get the packaging capacity.

10. Training is compression

  • Story: Ben's analogy — a 3 GB multi-layer Photoshop file (PSD) isn't something you send to a client; you compress it into a JPEG. The JPEG is more useful in most situations, but you can never get from the JPEG back to the PSD.
  • Insight: an LLM compresses the entire internet of text into a much smaller set of model weights; train once, infer endlessly, with inference relatively very cheap — the price is that redoing training is very expensive, so "you better do it right the first time."
  • Effect: the astronomical CapEx concentrates on the training side, which is exactly where Nvidia's revenue is; and the model spares everyone from remaking the PSD every time.

11. DGX Cloud: launch your own cloud through your rivals' clouds to reclaim the customer relationship

  • Story: CFO Colette Kress disclosed that about half of data-center revenue comes from CSPs concentrated in 5–8 companies — Nvidia owns the developer relationship through CUDA, but the customer relationship is intermediated by the clouds. DGX Cloud puts virtualized DGX systems inside Azure/Oracle/Google data centers, yet the customer contracts directly with Nvidia and logs in through Nvidia's website, starting at $37,000/month.
  • Insight: converting intermediated revenue into a direct sales relationship is worth more long-term than the accounting profit; the math is easy — building an equivalent A100 DGX costs ~$120,000, so $37,000/month is roughly a three-month CapEx payback for Nvidia and its cloud partner together.
  • Effect: David: "Nvidia launched their own cloud service through other clouds. This is unbelievable." Expect more fully owned-and-operated Nvidia data centers over time.

12. CUDA isn't moat rhetoric — it's the enabling layer that "gets the most people running on our hardware"

  • Story: since 2006 every card shipped can run CUDA; the developer curve: four years for the first 100,000, one million "13 years in" by 2016 (the transcript's phrasing, slightly off on the year, recorded as said), ~2 million around 2018, 3 million in 2022, and 4 million registered developers by May 2023.
  • Insight: insiders don't say "CUDA is our moat"; they say: to get people running on our hardware, we have to make it as easy as possible — so they carry 1,000–2,000 full-time software engineers building the compiler, runtime, debugger, CUDA C++, and industry libraries, while the software makes essentially no money (de minimis). "That is how you build a developer ecosystem."
  • Effect: taking the integral of headcount over the years, Ben estimates ~10,000 person-years poured into CUDA — "Good luck." — with 500 million CUDA-capable GPUs installed as a target base, and any existing AI code "just going to come right over and run within your brand new, shiny AI supercomputer."

13. The TAM narrative shifts from "a percent of every industry" to "the data-center replacement cycle"

  • Story: 2021 was the $100T × 1% method — Ben: "the topiest-down way I can think of to size a market." In 2023 Jensen reframed it: $1T of installed hard assets, $250B a year in CapEx, and Nvidia has "the most cohesive, fulsome, and coherent platform to be the future of what those data centers are going to look like."
  • Insight: Ben's "what do I have to believe" frame — you only have to believe one thing: AI workloads are creating real user value. The evidence: OpenAI is rumored to be over a $1B run rate (the most credible estimate Ben has heard is $3B, possibly a forecast for next year); ChatGPT is this boom's Netscape Navigator; the Fortune 500 wrote $10B of real checks last quarter. The one unknown: whether GPT-style experiences endure — "That is the thing you have to believe."
  • Effect: from hand-waving to a countable, verifiable pool of purchasing budgets; "the hype is actually showing up in revenue."

14. Do only what others can't; strike when the timing is right; ship everything you can during the window

  • Story: Intel used its "integrate into the motherboard → commoditize → build it ourselves" playbook against Nvidia, and later controlled the data-center interconnect via PCI Express, forcing Nvidia to "live in there… and I'm sure they hated every single minute of it" — an interviewee's oracle: "Intel was the country club and Nvidia is the fight club." Nvidia held off on a CPU for a decade until there was a real reason to differentiate, then shipped Grace.
  • Insight: don't chase low-margin opportunities; and a supply shortage isn't a moat but a window to be cashed in — data-center architecture decisions change every 5–10 years, so locking in now locks in for a decade, which is why you "ship as much product as you can while you have the lead" (David's add: happy to trade margin for throughput).
  • Effect: since Covid, back to a six-month ship cycle and two GTCs a year most years — "Imagine Apple doing two WWDCs a year." The de-slotting path is clear: plug into someone's server → build your own server → whole rows and walls → maybe run your own buildings, "we don't have to plug into anything."

Moat Analysis (the 7 Powers framework)

7 Powers is Hamilton Helmer's strategy framework (7 Powers: The Foundations of Business Strategy, 2016). This episode has no formal Grading segment — the hosts explicitly say the goal is "a more lasting big-picture" episode rather than a quarterly check-in, and the Bull & Bear cases take its place, noted faithfully.

The competitive landscape (Ben): Nvidia has a direct competitor (AMD), but that's not the most interesting form of competition — disintermediation is. AMD has neither TSMC's reserved 2.5D packaging capacity nor CUDA's developer ecosystem; the real competitive vectors are Amazon's homegrown Trainium/Inferentia, Microsoft's rumored AMD-partnered chip, Google's TPU, and Facebook using its PyTorch foothold to extend downward — plus data-center hardware vendors like Intel now becoming direct competitors. The core question: "Will Nvidia be able to stay ahead in ways that matter?" (Can it outrun a pack of chasers who think its margins are too fat and juicy and simply copy it?) The companion cloud discussion: Cloud 1.0's data lock-in (Snowball, literally driving hard drives to the cloud in semi trucks) means "you want to train your AI models right next to where your data is," but for the first time in five years Ben cocks his head at the cloud providers' moat — customers want "the full-stack Nvidia experience," not a cloud's cheapest-COGS approximation of it; and the very existence of billion-dollar GPU-dedicated clouds (Crusoe, CoreWeave, Lambda Labs) is smart money betting there's "a major cloud-sized opportunity." The convergent read: where you landed in Cloud 1.0 will strongly dictate where you land in the AI-cloud era — if customers demand Nvidia, the clouds have every incentive to serve it.

PowerVerdictEvidence
Counter-positioningNoneThe conclusion after the two debate it: nobody is actively choosing not to do what Nvidia does — everyone wants to be Nvidia right now. David floated "the existing data-center providers had incentives not to follow," but Ben stopped him with "what are they not doing now?" (they're all cramming GPUs into data centers) and David conceded: "Fair enough." The state of play is everyone chasing Nvidia's roadmap precisely, years behind
Scale Economies★ Strong"This has CUDA written all over it." With 4 million developers to amortize across, any fixed-cost investment is justified; 1,600 people on LinkedIn today have "CUDA" in their Nvidia job title (surely more); thousands of people on a software layer that makes essentially no money
Switching Costs★ StrongEverything of consequence (especially LLM training) is built on Nvidia — code plus organizational momentum; and bigger: NZS Capital's point that data-center revenue and CapEx is "some of the stickiest revenue that is known to humankind" — procurement lasts at least 5 years, and the architecture is standardized once a decade at most; even in a bubbly moment Nvidia is leveraging the excitement to buy lock-in
Cornered ResourceTextbook caseTSMC's CoWoS 2.5D packaging capacity, which competitors can't get (with a lucky element — some was reserved for crypto); Mellanox as the only InfiniBand-spec provider left. Ben: "It's like an invention delivered by aliens that very few humans know how to actually do." Important caveat: LLM training is a two-horse race — Google's TPUs are also produced at volume, but only through GCP, not an industry standard
Network EconomiesYes (compounds with Scale)Developers build libraries on CUDA and later devs call ready-made blocks — "You can write amazing CUDA programs that just don't have that many lines of code"; 500 million installed CUDA-capable GPUs are a target base. Ben self-corrects: probably more a scale economy; David cites Helmer and Chenyi — platform companies enjoy a special power that combines scale and network
Process PowerProbably weakest, but presentPro (David): the culture plus the six-month ship cycle would be very hard for competitors to match. Con (Ben): does it matter? Most workloads run fine on A100s; AMD already has 3D packaging on one of its latest GPUs (real copper-to-copper, no silicon interposer) more advanced than the H100's 2.5D — but "Nobody's going to make a purchase decision on this thing because it's a little bit of a better mousetrap." What matters is the system, the software, the ecosystem
BrandingYes, and right now very importantBen: "Nvidia is the modern IBM in the AI era — nobody gets fired for buying IBM"; decades of consumer brand carried into an enterprise posture, and even McDonald's CIO knows Nvidia; strength-leads-to-strength — last quarter's earnings is itself the strongest ad ("everyone else is buying Nvidia. I'd be an idiot not to"). David's close: "Nobody is getting fired for buying Nvidia anytime soon."

Gross margin and the "moat–castle relativity": today's 70%+ has a temporary component — the "blank check, I just need access to Nvidia hardware" scramble by enterprises and even governments (the UK, some Middle Eastern countries) will pass; but Ben judges the 65%+ level won't erode much. David's basis: to train a GPT-class model "there's one option" (you do it on Nvidia); inference is a more open market, but Nvidia is still best, and not for any single reason — "They're the best because of all of those" (hardware, data-center solutions, and CUDA — you need all three). The long-run sobering thought is Ben's castle metaphor: "Every moat only works if the castle is sufficiently small" — when the prize is a trillion-dollar market, this moat probably isn't wide enough, meaning margins come down and competition intensifies over time; Nvidia knows it completely, which is why it must be "pedal to the floor right now to outrun competition."

Bear Case (6)

  1. Whole-ecosystem siege: everyone in the tech ecosystem is now incentivized to take a piece of Nvidia's pie, with untold resources — witness Meta spending tens of billions on the metaverse, Apple a rumored $15B on its headset, Amazon tens of billions on devices (Echo never paid back); if the market really reaches $100B/year, at some point the game of chicken ends and someone goes all in. Hedge (David): "Never underestimate the inability of big tech to execute on stuff that it thinks it can. Especially with major strategy shifts."
  2. PyTorch aggregation: developers cluster on PyTorch → by Ben Thompson's aggregation theory, whoever aggregates customers can take margin and disintermediate; it has moved from Meta into a foundation that many companies contribute to; the CSPs' pitch will be "if you're building for PyTorch, it runs really well on our thing, too." (Ben: PyTorch vs Nvidia is an absolute false equivalence, but the direction is real; it still requires rewriting a lot of software underneath and shipping a lot of hardware.)
  3. The market isn't as big as the market cap reflects / a crisis of confidence: over the next 12–18 months there's a fair chance of a "maybe GPTs aren't as useful as we thought" awakening, a mini-bubble burst trickling to CIOs and CEOs — "it's not an if, it's a when." Hedge one: the top researchers said nearly in unison, "yeah, this is overhyped right now. Of course, obviously. But on a 10-year timescale you haven't seen anything yet." Hedge two: the hype is showing up as revenue — whoever's buying the compute believes something, and for Nvidia that belief is real dollars.
  4. Is generative AI all it's cracked up to be: David personally still can't use ChatGPT (hallucination, doesn't fit his workflow; a month earlier he was "pounding the table" about it), but concedes Acquired is a "hyper specialized, unique little unicorn" and not representative. The counter-examples are strong: GitHub Copilot gives leverage at every developer level from zero to elite (Ben is a firm believer); Runway (CEO Chris Valenzuela, usually rendered Cristóbal Valenzuela) technology was used in Everything Everywhere All at Once, "just the tip of the iceberg"; some people use ChatGPT to set their OKRs. Ben's frame: AI's value is the sum of countless niches — you just happen not to be in one of the first few. There's also the "timing too perfect" suspicion: VCs pitched crypto as the future, rates hit 5%, and now they pitch AI (David: "come on guys. It's too perfect.") — but the latest quarter proves the Fortune 500 is really buying.
  5. The training → inference shift: the popular narrative is that models eventually get "good enough, all trained," and the load shifts to less-differentiated inference. Ben isn't fully convinced: the Transformer isn't the end of the road, and things beyond it are already in research; but there's a real second bear layer — this round threw a "brute force kitchen sink" at training, and all that revenue accrued to Nvidia (the kitchen-sink maker), while Google's Chinchilla and Llama 2 approach GPT-4 quality with fewer parameters (judge for yourself), so smarter models may need less compute. Besides, most AI workloads don't look like LLMs anyway (the whole diffusion / image-generation genre is far less compute-intensive) — "LLMs are the current maxima in human history of jobs to be done that require a ton of compute." David partly agrees: training won't vanish, but inference's share grows, and there the ecosystem really is less differentiated.
  6. China: a legitimate and real bear — 25% of last fiscal year's revenue, unaddressable in a meaningful way for the foreseeable future; China is racing to build its own homegrown ecosystem and competitors, and the market will close off — "What's going to come out of there?" (an open-ended risk).

(The seventh candidate — "high-multiple collapse" — the two flip into a bull: a post-spike stall is fatal for most companies via stock comp, morale, and customer perception; but Nvidia has repeatedly risen from the ashes after years of terrible sentiment with mind-blowing innovation, making it "the company with the best disposition to handle that when it happens.")

Bull Case (5)

  1. Jensen is right about accelerated computing: only ~5%–10% of workloads are accelerated today, heading to 50%+, so total parallel compute explodes and mostly accrues to Nvidia. David's nuance: it isn't that traditional compute disappears (the SharePoint servers keep running) — AI compute is added on top of everything. Ben's engineering view: parallel code is really hard (race conditions and semaphores are the hardest things to debug in a CS class), and many workloads are unaccelerated only because it's hard to develop for; if Nvidia's full stack makes migration easy enough, "latent accelerated addressable computing" gets released in batches.
  2. Jensen is right about generative AI: ChatGPT is rumored to be over a $1B run rate (most credible estimate $3B); Google Bard isn't directly monetized but kept prep-mode Ben on Google search. "The bear case is that everything has to go right for Nvidia; the bull case is indications that things are going right for Nvidia."
  3. Nvidia just moves fast: whatever develops technically, it's hard to believe they won't find a way to be well-positioned — a purely cultural factor.
  4. The data-center replacement market: $1T installed, $250B/year spent; Nvidia's data-center business annualizes to ~$40B (a single quarter of $10.3B), roughly 15%–18% of current data-center spend (Ben first misspoke 20%, David corrected by stripping out gaming) — room in both the share and the denominator. Cross-check via supplier CapEx: TSMC's latest earnings say AI hardware is only 6% of revenue, but they expect AI revenue to grow 50% a year for five years and are building new fabs and packaging with real money — "it's expensive for TSMC to be wrong," and suppliers voting with dollars is more credible than narrative.
  5. Nvidia isn't Intel, isn't Cisco — it's Microsoft, maybe even old-school IBM: the popular "we've seen this movie — a hardware cycle stock" misses that Cisco has no developers and Intel never did ("Microsoft had developers and Intel had Microsoft"); Nvidia has developers — it built a non-von-Neumann computer with a whole new language/compiler/computing model: "That's CUDA and it fricking works"; the self-image is "We are a foundational computer science company. We're not just slinging hardware here." David goes further: it's old-school IBM — IBM ruled the B2B mainframe cycle before the PC (today's Edge) disrupted it; if Jensen is right, we're returning to a "centralized data center = modern mainframe" cycle, this time in a computing market orders of magnitude larger. Ben: "that's the biggest realization that you helped me have." (Consensus footnote: a lot of inference will move to the Edge, but training won't anytime soon.)

The Nine-Level Thought Experiment: "what it would take to compete with Nvidia head-on" (recorded in full)

Ben's closing walkthrough, each step joined by "even if you did that":

  1. Design GPU chips that are just as good — AMD, Google, and Amazon are arguably doing this;
  2. Build chip-to-chip networking like NVLink — very few have it;
  3. Build relationships with assemblers like Foxconn to turn chips into DGX-class servers;
  4. Build server-to-server and rack-to-rack networking as good as Mellanox's best-on-the-market InfiniBand — which Nvidia now fully owns, so basically nobody has it;
  5. Convince customers to buy your thing — it must be better or cheaper or both, and (David's add) better by a wide margin, because "you're not going to get fired for buying Nvidia";
  6. Contract TSMC's most advanced 2.5D CoWoS packaging capacity — which is gone, so good luck;
  7. Build software as good as or better than CUDA — ~10,000 person-years, costing not just billions of dollars but the time itself;
  8. Convince developers to drop CUDA for you;
  9. And Nvidia won't be standing still — you have to do all of the above in record time and also catch whatever it has newly developed in the meantime.

Conclusion: head-on competition is nearly impossible; the only things that could unseat Nvidia are an unknown flank attack it doesn't see, or a future that simply isn't accelerated computing and AI — which seems very unlikely. Coda: back in 2015–2016 on Acquired, Marc Andreessen already noted that A16Z was seeing every deep-learning startup built on Nvidia — in hindsight, rather than fund all the AI startups, they should have put every dollar of every fund into Nvidia at the market price. David: "Mark is right once again. Strength-leads-to-strength."

Deep Cuts

  • The Rosewood dinner: 2015, Musk and Altman host nearly all the top researchers under the duopoly on Sand Hill Road to lure them out. On the scene (David): what would it take for you to leave? "The answer from almost all of them is nothing… We're happy as clams here." Ben's kicker: "There's a money spigot pointed at our face" — and, more importantly, the chance to work with all the best people in the field. Cade Metz's record: no one was sure these minds could be pulled to a startup even with Musk and Altman behind it, "but one key player was at least open to the idea of jumping ship" — Ilya.
  • Microsoft's three rounds: $1B in 2019 (for exclusive cloud provider status) → $2B in 2021 → $10B in Jan 2023 (GPT into all products); plus a September 2020 exclusive commercial license to the underlying model. What it means for Nvidia: OpenAI's compute is all bought through Azure, and Azure runs Nvidia.
  • DGX bundle economics: eight $40,000 H100s = $320,000, but a full DGX H100 sells starting at $500,000 — a ~$180,000 spread, with that ARM CPU costing them "essentially nothing" to make. David's recurring bit: "you say solution, I hear gross margin." Ben: this is bundle economics — bundle more value, earn more margin; but right now the margin isn't coming from "solution" at all, purely from supply shortage — "I'll write you a blank check and Nvidia you write whatever you want on the check."
  • Cloud rental reality: an 8×A100 DGX server lists at ~$30/hour (Azure/AWS); AWS's P5.48xlarge (8×H100) is ~$100/hour — "when I say you can get access, I don't actually mean you can get access. That's the price." The DGX GH200 SuperPOD (a 256-rack "AI wall," the first turnkey AI data center, able to train a trillion-parameter GPT-4-class model) is priced "call us" (Ben: "Of course it is"; David guesses hundreds of millions).
  • The H100 baseball card: ~a quarter-trillion transistors, 35,000 parts, 70 pounds; 18,500 CUDA cores, 640 Tensor Cores (specialized for matrix multiplication), 80 streaming multiprocessors; 30x faster overall than the just-2.5-years-older A100, and 9x faster for AI training; it takes robots to assemble and AI to design. Ben: "They have completely reinvented the notion of what a computer is."
  • The Megatron–Mellanox timeline easter egg: the acquisition was announced (March 2019) five months before Megatron shipped (August 2019) — internal research ahead of public evidence; Ben found this out two hours before recording.
  • The write-down became a blessing: when crypto crashed in 2022, Nvidia wrote down its pre-ordered TSMC capacity as a "really big blemish" — and now "oh my God, are they glad that they reserved all that capacity."
  • The "disappointing year" paradox: Jensen told Stratechery in March 2023 that "last year was unquestionably a disappointing year" — the year ChatGPT launched. The market-cap rollercoaster proves it: 8th largest in the world (~$660B) in April 2022 → below $300B → back over $1T within months.
  • Naming: Grace CPU + Hopper GPU = Grace Hopper (the great computer scientist and US Navy Rear Admiral); the consumer Lovelace architecture honors Ada Lovelace.
  • Headcount efficiency: Nvidia has 26,000 employees; Microsoft's market cap is only 2x but it employs 220,000 — ~$46M of market cap per Nvidia employee (Ben notes the figure is a bit farcical, since the huge market cap is recent).
  • Has Google X ever shipped anything profitable: Ben asks. David: "Google Brain." — "We'll leave it at that." Astro Teller's line to the NYT: the profit gains from the Google Brain team alone more than funded everything Google X was doing — and that's not even counting DeepMind.

Era & Industry Trivia (tangents worth keeping)

  • Hinton was born a Boole: Geoff Hinton is the great-great-grandson of George and Mary Boole — AND, OR, XOR, NOR, the whole Boolean-algebra family, all come from that couple. Ben: "I also didn't know there were people named Boole, that that's where that came from. That's hilarious."
  • ImageNet's crowdsourced base: 14 million hand-labeled images, the largest use of Mechanical Turk up to that point, with labelers paid ~$2 an hour — and at internet scale that dataset is still "a drop in the bucket."
  • The duopoly era's spoils: YouTube went from a money-losing "crazy acquisition" to a juggernaut — the feed and autoplay all came out of AI research (before that, most views were embeds on other web pages); DeepMind cut Google's data-center cooling costs; Instagram became a $100B–$500B asset for Meta via AI recommendation feeds (Ben self-evidences that targeted ads work: "I have bought a lot of things on Instagram ads"). The counterpoint: outside the duopoly, AI was bad — Siri was terrible; David's straw man: Snap and Musical.ly (sold to ByteDance, becoming TikTok) never reached independent scale, maybe because they couldn't get those researchers (Ben: "a couple of steps too far… but still a fun straw man").
  • The TPU/TensorFlow soft-hard gambit: open-source framework → the framework runs best on the hardware optimized for it → the hardware is only on Google Cloud. People gasped, "why is Google giving away the farm for free?" — it was 3–4 years early and prescient. But the TPU became a casualty of a strategy conflict: it's one of LLM training's "two horses," yet only available through GCP, never AWS/Azure — while zooming out, it makes sense for Google: the TPU is mostly consumed internally (Bard, generative AI in search), and the core business is one of the most profitable cash geysers ever, so anything that extends its runway is worth doing.
  • OpenAI's prehistoric projects: during the Transformer window, still doing "researchy" computer vision — the DOTA II bot (beating the world's best purely from screenshots; David: "a faster horse… for the past generation") and Universe (using the Grand Theft Auto world to train self-driving vision — "crazy stuff, but it was scattershot," nobody mentions it now). The defense: they knew everything about the Transformer, they just couldn't afford the training bill as a nonprofit.
  • Export controls prove the size of the lead: September 2022 Biden-administration export controls on advanced computing to China (Ben corrects David: controls, not bans); the A800/H800 are nerfed SKUs with cranked-down NVLink transfer speeds, and they're still "selling like hotcakes" in China — "still the best hardware and platform that you can get in China, even a crippled version" (David: "I can't think of a better illustration of just how wide their lead is"); Chinese companies are stockpiling A800s anticipating tighter controls; mainland China was 25% of last fiscal year's revenue, mostly cloud players like Baidu, Alibaba, and Tencent; Baidu's GPT competitor has over a trillion parameters and may be larger than GPT-4.
  • The InfiniBand survivor story: it was an open consortium standard, far faster than ethernet, but the consensus that "ethernet is the lowest common denominator — everyone had to implement it anyway" pushed vendors out, leaving Mellanox alone — until the era of training giant models across racks made 3,200 gigabits/second suddenly the crux.
  • Ethereum's parting gift: the move to proof of stake ended GPU mining demand, directly driving Nvidia's 2022 inventory write-down — which indirectly left behind the later-cherished TSMC reserved capacity.
  • The post-von-Neumann frontier: academia is experimenting with compute-in-memory — rather than moving data over copper to the CPU (lossy, expensive, energy-intensive), process it right where it sits in memory. "They really are rethinking what is a computer?"
  • Governments are queuing too: the posture of the UK and some Middle Eastern countries — "blank check. I just need access to Nvidia hardware."
  • The Echo paradox: David standardized his whole house on Amazon Echo and is indignant: "How in this world of incredibly accelerating AI capabilities are my Echoes getting dumber?" Ben's zinger: "They need to train them in Inferentia a little bit harder." David: "Oh, Jesus. Okay, rant over."
  • The too-smooth VC pivot: "you just told everybody about how crypto's the future… then interest rates went to 5%… and now the future is AI. This is the best time ever to be investing" — David: "come on guys. It's too perfect."

Cross-domain Notes

This is the Acquired episode with the strongest overlap with the PH argument network, and the direction is unmistakable: this is independent evidence from the business domain — neither host builds a geopolitical narrative; they reason strictly from commercial and technical logic, yet converge with several core PH judgments, and the evidentiary value lies precisely in that independence. The 2023-09 recording date also makes it a clean time-anchor.

  • The power structure of AI compute → ai-power-structure: this episode sketches a textbook sample of power concentration — top talent first locked up by the Google/Facebook duopoly ("a money spigot pointed at our face"); the very organization built to break that monopoly, OpenAI, is itself "powers on high and existing money… will[ing] something into existence" (Ben's words), and is then pushed into Microsoft's arms by compute costs; half of data-center revenue concentrated in 5–8 CSPs; "there's one option" for training a GPT-class model; the UK and Middle Eastern governments queuing with blank checks. The structure "whoever controls the training infrastructure controls the spectrum of AI capability" is here supported entirely by business-side data.
  • Chip export controls → PH's technology-war front: the September 2022 China controls, the nerfed A800/H800, the 25% China revenue exposure, stockpiling ahead of tightening, and Baidu's trillion-parameter model — the chip-war narrative reflected directly at the level of corporate earnings; the detail that "the crippled version is still the best platform China can buy" simultaneously quantifies the size of the US technical lead and the practical limits of the controls, usable as time-anchored evidence for US-China tech-decoupling arguments.
  • The data-center empire and the centralized-computing cycle → technate / ai-apocalypse: David's "back to the mainframe cycle" judgment (centralized data center = modern mainframe, Nvidia = the IBM of that cycle), Jensen's "the data center is the computer" and the de-slotting endgame (owning and operating whole buildings), and the $1T installed-base replacement narrative — all give the "concentration of technocratic infrastructure" thesis under technate an industrial-side skeleton: energy, compute, models, and cloud are verticalizing into a very small number of nodes. And "training is compression" (the whole internet packed into unrecoverable weights), scaling's emergence that "no researchers expected," and the top researchers' collective "on a 10-year timescale you haven't seen anything yet" are an in-industry statement of the ai-apocalypse proposition that capability is growing faster than understanding and governance.
  • Methodological resonance: Ben's "what do I have to believe" frame and the "validate the narrative against supplier CapEx" move (TSMC voting with new fabs) are structurally identical to the PH domain's practice of calibrating predictions against real-money behavior rather than words, and can be cited as a forecasting-calibration tool.

Pages Worth Creating

  • Entities: Jensen Huang(黄仁勋) (founder page; a connector to the Trader Joe's episode via the Denny's coworker anecdote), TSMC:纯代工模式发明者 (a dedicated TSMC episode page; CoWoS / Cornered Resource cross directly with this episode), openai (the episode's real second protagonist: the full structural arc of nonprofit → capped-profit → Microsoft alliance)
  • Concepts: 7 Powers 护城河框架 (Hamilton Helmer's framework, the standard analytical toolkit across the Acquired series), cuda (the textbook case of a developer-ecosystem moat: ~10,000 person-years, 4 million developers, 500 million installed GPUs)
  • Cross-domain: ai-power-structure and technate can both cite this episode as an independent business-domain evidence node

Source · acquired