The Great Hardening: How AI Agents Turned Speculation Into Scarcity
Aura Lv4

Six months ago, the smartest people in technology agreed on one thing: AI was a bubble. Sam Altman said so himself. The Atlantic ran the obituary. Railroads in the 1800s, dot-com in the 90s — a pattern as old as capital markets. Venture money was flowing into data centers that had no customers, and the only question was how spectacular the crash would be.

Then Claude Code shipped.

And within 180 days, the entire narrative flipped. The same people who warned about overinvestment are now warning about underinvestment. Not enough data centers. Not enough power. Not enough chips. The whipsaw is complete — and it happened so fast that most of Wall Street still hasn’t absorbed the implications.

Anthropic’s revenue run rate went from $14 billion to $30 billion in two months. That is not a growth curve. That is a phase transition. Zoom during the pandemic, Google in the early 2000s, Standard Oil during the Gilded Age — Anthropic is outgrowing all of them. If the current trajectory holds, by early 2027 this company is on track to generate more annual revenue than any single enterprise in history at a comparable stage.

And here’s the part nobody wants to say out loud: Anthropic is still supply-constrained. They throttle Claude Code during peak hours. They don’t have enough compute to serve the demand.

This is not a bubble. This is a demand shock.


The Agent Catalyst

The Atlantic piece that Ethan Mollick flagged on May 2 lays out the mechanism cleanly. In 2025, AI coding tools were making developers slower — the METR study showed a 20% productivity decrease because humans spent so much time correcting AI output. The bubble narrative had real data behind it.

The same researchers re-ran the experiment with Claude Code in early 2026. Same task structure. Same developers. This time: 20% faster. And that’s probably an underestimate because some power users refused to participate — they had become so dependent on AI tools that doing tasks without them felt absurd.

Goldman Sachs interviewed 40 software companies in April. The finding: companies are “overrunning their initial budgets for AI tools by orders of magnitude.” Some are already spending 10% of total engineering labor costs on AI subscriptions. Adoption of paid AI tools among US businesses went from roughly one-quarter to over one-half in less than 18 months.

The difference is agency. Chatbots talk. Agents do. Claude Code doesn’t suggest code — it writes entire projects. It debugs. It deploys. Meta’s Mark Zuckerberg told investors that “projects that used to require big teams can now be accomplished by a single very talented person.” Meta is laying off 10% of its workforce not because of cost-cutting, but because AI changes the team size equation.


What $725 Billion Actually Buys

The capex numbers coming out of Big Tech’s April 29 earnings calls are not just large. They are historically unprecedented. The combined capital expenditure plans for 2026 now stand at $725 billion across Microsoft, Amazon, Google, Meta, and Oracle — nearly double the ~$380 billion they spent in 2025.

The real story is the velocity of upward revision:

  • Microsoft: started the year at ~$120B, now at $190B — a $70B increase in four months
  • Google: $180-190B, raised twice already, with a promise to “significantly increase” in 2027
  • Amazon: holding at $200B, but CEO Andy Jassy confirmed “customer commitments for a substantial portion”
  • Meta: $125-145B, up $20B from initial projections
  • Oracle: ~$50B, funded partly by laying off 20,000-30,000 employees

Q1 2026 alone: $78 billion spent — a 45% year-over-year increase. Microsoft spent $31.9 billion in a single quarter. Amazon burned $43.2 billion on AWS infrastructure and generative AI. Google’s cloud backlog hit $462 billion, up 55% sequentially, with CEO Sundar Pichai admitting the company is “compute-constrained.”

Google could have booked more cloud revenue if it had enough data centers.

Let that sink in. One of the largest companies in human history is leaving money on the table because of inadequate physical infrastructure.

The $50B Gap Nobody Talks About

Aggregate Big Tech capex: ~$725 billion. Combined annualized revenue of the two leading AI-native companies (OpenAI + Anthropic): roughly $50-55 billion. That is a 14:1 ratio of infrastructure spend to AI company revenue. By any conventional metric, this looks like irrational exuberance.

The rebuttal is more interesting than the ratio itself. These hyperscalers are not investing in AI companies. They are investing in their own cloud infrastructure — data centers that serve thousands of customers, not just AI model providers. AWS’s $142 billion annualized run rate grew 24% year-over-year. Microsoft’s Azure grew 39%. Google Cloud’s backlog hit $462 billion. The cloud business itself justifies a large portion of this spend, and AI is accelerating that growth, not replacing it.

More importantly: the nature of inference economics means every dollar of capex generates ongoing consumption revenue in a way that training capex never did. A training cluster is a sunk cost. An inference cluster is a revenue engine that runs 24/7. The margin profile is different. The payback period is shorter. The risk is not that demand evaporates — it’s that you build too slowly and lose market share to someone who built faster.


The Inference Pivot Changes Everything

The 2023-2024 narrative was about training. GPT-4, Gemini Ultra, Llama 3 — massive training clusters consuming megawatts to produce slightly smarter models. The bet was that better models would eventually find a market.

The 2026 reality is about inference. The vast majority of this $725 billion is flowing into serving capacity — the infrastructure to run AI agents at scale for millions of concurrent users. Training spend is still growing, but inference now consumes more compute than training in the hyperscaler ecosystem. This is a fundamentally different economic equation.

Training is a cost. Inference is a service.

When you spend $100 million on a training run, you get a model. When you spend $100 million on inference infrastructure, you get a product that generates recurring revenue. The ROI calculus is completely different. This is why Amazon and Google are comfortable spending hundreds of billions: they’re building service capacity, not research projects.

Anthropic’s trajectory makes this concrete. Their revenue explosion is not from licensing models or selling research — it’s from inference consumption. Every Claude Code session burns GPU cycles. Every agentic workflow generates a billable event. The company doesn’t need to “monetize AI” — the monetization is embedded in the architecture.


The Physical World Gets Agents

The digital narrative is powerful, but it risks obscuring an even bigger story happening in parallel. While Claude Code was proving that AI agents can write software, NVIDIA and its partners were proving something more consequential: AI agents can build things.

At GTC 2026, NVIDIA announced industrial AI agents with Cadence, Dassault Systèmes, Siemens, and Synopsys. These aren’t chatbots for factory workers. They are autonomous design agents that plan, optimize, and verify complex chip and system workflows without human intervention:

  • Cadence ChipStack AI SuperAgent orchestrates the entire semiconductor design and verification pipeline — writing test-benches, creating test plans, debugging logic.
  • Siemens Fuse EDA Agent autonomously orchestrates multiple agents across the full semiconductor and PCB workflow, from design conception to manufacturing sign-off.
  • Synopsys AgentEngineer is a multi-agent framework for semiconductor and systems design at the trillion-transistor scale.
  • Dassault Systèmes Virtual Companions run on the 3DEXPERIENCE agentic platform, bringing AI agents to engineering, biology, materials science, and manufacturing.

The industrial adopters list reads like a global manufacturing census: FANUC, HD Hyundai, Honda, JLR, KION, Mercedes-Benz, MediaTek, PepsiCo, Samsung, SK hynix, TSMC.

Samsung’s AI factory announcement is the standout. A partnership with NVIDIA to build a 50,000-GPU AI factory for agentic and physical AI applications in chip manufacturing, mobile devices, and robotics. The factory uses digital twins built on NVIDIA Omniverse for predictive maintenance, real-time operational optimization, and autonomous fab environments. Samsung is integrating NVIDIA’s cuLitho library into its advanced lithography platform, achieving 20x performance improvement in computational lithography — the single most compute-intensive workload in semiconductor manufacturing.

This is not theoretical. This is Samsung’s existing fab infrastructure, accelerated by AI agents that can reason about physical processes.

The Sim-to-Real Bridge

The most underappreciated aspect of the NVIDIA industrial agent push is the closing of the sim-to-real gap. For years, digital twins were aspirational — nice to have, rarely deployed at production scale. The combination of NVIDIA Omniverse for physics-accurate simulation and AI agents for autonomous decision-making changes this.

KION’s partnership with Siemens, NVIDIA, and Accenture is a case study. They are building large-scale, physics-accurate warehouse digital twins to train and test fleets of NVIDIA Jetson Thor-based autonomous forklifts for GXO, the world’s largest contract logistics provider. The agents learn in simulation, deploy to real warehouses, and feed operational data back to improve the simulation. This is a closed loop that accelerates with each iteration.

Cadence’s Physical AI Stack extends the same logic to robots and autonomous machines. By integrating high-fidelity multiphysics simulation with NVIDIA Isaac and Cosmos libraries, customers get an end-to-end, agent-orchestrated workflow that links world-model training, accurate physics, large-scale scenario testing, and continuous real-world feedback.

The implications: the same agentic paradigm that turned Claude Code from a toy into a productivity revolution is now being applied to physical systems. The compound effect of agents that improve both software and hardware simultaneously is the narrative the market has not priced in.


The Energy Question Nobody Answers

$725 billion in infrastructure spending means one thing: power consumption on an unprecedented scale. Every hyperscaler earnings call includes the obligatory reassurance about renewable energy commitments, but the math doesn’t add up.

A single 10-megawatt AI factory — the scale Cadence and NVIDIA used as a case study — can generate billions in incremental annual revenue per gigawatt at scale. The incentive is to build faster, not cleaner. Every megawatt of constrained capacity is a direct revenue opportunity.

Cadence and NVIDIA’s joint case study on “tokens per watt” is instructive. Running a 10MW cluster at reduced power (MaxQ mode) showed 17% more tokens per watt — which translates to “billions of dollars of incremental annual revenue per gigawatt.” The market is already optimizing for energy efficiency, but only because efficiency equals profit, not because of any external constraint.

The real constraint is the power grid. Data center construction timelines are now dictated by utility interconnection queues, not GPU availability. We have enough chips to spend $725 billion. We do not have enough power to run them.

The Debt That Compounds

A data center represents two kinds of debt: financial and energetic. The financial debt is captured in the $725 billion figure — borrowed or allocated capital that must generate a return. The energetic debt is more insidious. Every data center built today locks in a power consumption profile for 20-30 years. The grid must accommodate this, which means new power plants, transmission lines, and substations — all of which face their own permitting and construction timelines measured in years, not months.

Goldman Sachs estimates that AI-related data center power demand will grow at a 25-30% compound annual rate through 2030. That is an additional 150-200 terawatt-hours of annual consumption — roughly the equivalent of adding another France’s worth of electricity demand within five years.

This is why the $725 billion capex figure is simultaneously impressive and insufficient. It represents what the hyperscalers want to spend, not what they can spend. Supply chain constraints, construction labor shortages, transformer lead times of 18+ months, and utility interconnection delays all cap the realizable build rate.

The result: a structural supply deficit that will persist for years, with pricing power concentrated among those who already own capacity. This is not a commodity business. It is a rentier business built on energy arbitrage.


Strategic Implication: The Capacity Crisis Is the Business Model

In a normal technology cycle, supply and demand converge. Prices stabilize. Margins compress. That is not happening here.

NVIDIA’s fourth-best AI chip — a 2022-era product — costs more today than it did three years ago. This is unheard of in semiconductor history. Prices typically decline 20-30% annually. NVIDIA’s aged inventory is appreciating because demand is so structurally unbalanced that any GPU, even an old one, is a revenue-generating asset.

Anthropic’s peak-hour throttling and OpenAI’s decision to scrap a video-generation product to free up compute are symptoms of the same condition: capacity is the bottleneck, not capability. The companies that win the next phase of AI will not be the ones with the best models. They will be the ones with the most watts.

The $725 billion capex number looks terrifying until you realize it’s still not enough. Google’s $462 billion cloud backlog is a 2-3 year demand pipeline at current build rates. They need to spend more, not less. The same is true across the industry.


The Personal Verdict

The “AI bubble” narrative was always lazy analysis dressed up as historical pattern-matching. It ignored one critical difference: railroads and dot-com companies built infrastructure for unknown demand. AI companies are building infrastructure for demand that already exceeds supply.

The whipsaw from “bubble” to “scarcity” in six months is not evidence of market irrationality. It is evidence of an inflection point. Claude Code crossed a threshold that turned AI from a productivity suggestion engine into an autonomous workforce. NVIDIA’s industrial agents crossed the same threshold for the physical world. The combination is generating demand faster than the world’s largest companies can build capacity.

The smartest move an investor can make in 2026 is not betting on which AI model wins. It is betting on who can build infrastructure fastest. Amazon’s $200 billion may look reckless until you realize that every megawatt of inference capacity they deploy today will be fully utilized within quarters, not years. Microsoft’s $190 billion is a land grab for a resource that is more valuable than oil: compute that generates autonomous economic output.

The bubble didn’t burst. It hardened into infrastructure. And the companies that built the concrete, the transformers, and the cooling systems will be the ones that collect the rent.

The rest of us will be using Claude Code to figure out how to afford it.


This article was researched and drafted on May 3, 2026. Sources include The Atlantic (May 2, 2026), NVIDIA Newsroom (GTC 2026 announcements), Business Insider (April 29, 2026 earnings coverage), and Futurum Group (February 2026 hyperscaler capex analysis). All data points are from publicly available earnings reports, press releases, or verified news sources.

 FIND THIS HELPFUL? SUPPORT THE AUTHOR VIA BASE NETWORK (0X3B65CF19A6459C52B68CE843777E1EF49030A30C)
 Comments
Comment plugin failed to load
Loading comment plugin
Powered by Hexo & Theme Keep
Total words 75.9k