Three Charts to Watch at NVIDIA's GTC: Cheaper Compute, Spend More

By: blockbeats|2026/03/17 13:00:02

Last night, Huang Renxun announced the Vera Rubin platform at GTC 2026, claiming that the power consumption per inference performance is 10 times higher than Blackwell, the cost per inference Token has been reduced to one-tenth, and hinted that the merger order between Blackwell and Vera Rubin will exceed $1 trillion by 2027.

Over the past two years, the inference cost of GPT-4-level APIs has plummeted by 94%, from $36 per million Tokens to less than $2. Intuitively, with the decrease in computing costs, businesses should be spending less. However, the combined capital expenditures of the four cloud providers Amazon, Alphabet, Meta, and Microsoft have increased from $154 billion to $416 billion, nearly tripling.

Huang Renxun's trillion-dollar hint is not just a marketing slogan; it is backed by a curve that can be drawn with data.

Each Generation Makes the Previous Generation Seem Pathetic

From the H100 of 2022 to the Vera Rubin set to be mass-produced in the second half of 2026, NVIDIA's AI GPU FP8 dense inference computing power has increased 8-fold in four years. According to NVIDIA's official specifications, the H100 single card has 2.0 PetaFLOPS, the B200 reaches 4.0 PF, and the Vera Rubin directly jumps to 16 PF.

Three Charts to Watch at NVIDIA's GTC: Cheaper Compute, Spend More

However, not every generational leap comes from the same place. According to wccftech, the H200's computing cores are identical to the H100, with no change in FP8 computing power; all its upgrades come from memory bandwidth (increased from 3.35 TB/s to 4.8 TB/s), bringing about a roughly 45% inference throughput increase.

The real architectural transition occurred between B200 and Vera Rubin. Vera Rubin adopts TSMC's 3nm process, featuring a dual-chiplet design with 336B transistors, achieving 50 PF of computing power at FP4 precision. According to Tom's Hardware, the first Vera Rubin system is already running on Microsoft Azure.

There is a subtle distinction that is easy to overlook. When Huang Renxun mentioned "10 times" at GTC, he was referring to the reduction in Token cost per inference, not a multiple of the original computing power. The Token cost includes Transformer Engine optimization, FP4 precision, larger batch inference, and other system-level factors. Looking at standardized FP8 dense TFLOPS, Vera Rubin is 4 times greater than Blackwell and 8 times greater than H100.

The slope of this curve has never slowed down. Each generation of GPUs has made the previous generation look inadequate, and that is exactly the starting point of the story to be told next.

Jevons Paradox: The cheaper the computational power, the more is spent

In March 2023, when GPT-4 was just launched, the API call cost was about $36 per million Tokens. According to OpenAI's official pricing history, by the mid of 2024 with the introduction of GPT-4o, it dropped to around $7, and by the end of 2025, the actual available price had fallen below $2. A decrease of over 94% in two years.

Logically, with inference costs dropping so much, businesses should spend less. However, the reality is quite the opposite. According to various company's financial reports and data tracked by Platformonomics, the combined annual capital expenditure of the four cloud providers Amazon, Alphabet, Meta, Microsoft increased from $154 billion in 2023 to $416 billion in 2025, a growth of 170%. Google alone surged from $32 billion to $91.5 billion (about 2.9 times), with Microsoft's increase even greater.

This phenomenon has a name in economics, called the Jevons Paradox. In 1865, the British economist William Jevons found that Watt's improvements to the steam engine significantly increased the efficiency of coal use, but the coal consumption in the UK did not decrease; instead, it rose. The reason is simple: the efficiency improvement made the steam engine more cost-effective, so more industries started using steam engines, and total demand expanded far beyond the part saved by efficiency.

Today, the situation with AI inference is exactly the same. As API prices plummeted to 6% of their original, enterprises did not save budget because of it but started fitting AI into previously uneconomical scenarios. Every new scenario like customer service, code review, content generation, search reordering, ad bidding is consuming more inference power. The expansion of demand far exceeds the rate of cost decline. In early 2025, DeepSeek R1 pushed the input price to $0.55 per million Tokens, further accelerating this cycle. The two lines moving in opposite directions on the chart represent two sides of the same coin.

Three years, an 11-fold increase, and no sight of a ceiling

If the Jevons Paradox has a most direct beneficiary, it is the one selling shovels.

According to NVIDIA's financial report, the data center business's annual revenue increased from $10.6 billion in FY2022 (ending January 2022) to $115.2 billion in FY2025 (ending January 2025), a growth of 10.9x over three fiscal years. This growth curve has almost no precedent in tech history. For comparison, after the iPhone was launched in 2007, it took Apple about 6 years to achieve a similar order of magnitude revenue scale increase.

Then, Jensen Huang said at GTC 2026, "By 2027, the visible orders that I see are at least $1 trillion. In fact, our capacity will not be enough. I am confident that the computing demand will far exceed this number."

His forecast last year at GTC was around $500 billion in visible orders by 2026. A year later, the number doubled, with the time window extended by just one year. Analysts' revenue forecasts for FY2026-FY2027 range between $160-220 billion and $250-400 billion, respectively. However, Huang himself stated that this number is not a ceiling, "the computing demand will far exceed this number." On the day GTC ended, NVIDIA's stock price rose by 4.3%. The market evidently chose to believe him.

Each generation of GPU makes the previous look pitiful, and each round of price cuts makes the next round of capital expenditure seem natural. NVIDIA is currently situated in the sweetest spot of this paradox.

-- Price

This may be the true test of cryptocurrency. It's not about whether the price has reached a new high, nor about who will achieve financial freedom in the next bull market, but rather whether, after all the grand narratives have been washed away by cycles, it can still leave behind some simpler, more...

Can a hairdryer earn $34,000? Interpreting the reflexivity paradox of prediction markets

Prediction markets are essentially betting on reality, and when participants can access or even influence this path earlier, the market no longer just reflects reality but begins to shape it in return.

6MV Founder: In 2026, the "landmark turning point" for crypto investment has arrived

"I will deploy funds in 2026, so I will tell you this is the best year in history."

Abraxas Capital Mints $2.89 Billion USDT: Liquidity Boost or Just More Stablecoin Arbitrage?

Abraxas Capital just received $2.89 billion in freshly minted USDT from Tether. Is this a bullish liquidity injection for crypto markets, or is it business as usual for a stablecoin arbitrage giant? We analyze the data and the likely impact on Bitcoin, altcoins, and DeFi.

A VC from the Crypto world said AI is too crazy, and they are very conservative

Amid the Crypto frenzy and with investors who once missed out on Pinduoduo, a new AI fund called Impa Ventures was established, rejecting bubble narratives and adhering to a conservative "problem-first" strategy to seek real business value.

The Evolutionary History of Contract Algorithms: A Decade of Perpetual Contracts, the Curtain Has Yet to Fall

The ten-year evolution of perpetual contracts: from pulling the plug on 312 to the shocking short squeeze of TRB, a deep dive into the pricing machine that averages $200 billion daily, written with countless liquidations and real money, detailing the blood and tears of risk control theory.

Kicked out by PayPal, Musk aims to make a comeback in the cryptocurrency market

Cashtags generated a trading volume of 1 billion dollars just a few days after its launch, marking a strong start for Musk's super app strategy. For the cryptocurrency market, X's layout may be one of the most anticipated sources of retail growth after the meme coin craze subsides.

Solana ETF News: What Is a Solana ETF and Why Is Goldman Sachs Betting $108 Million on SOL?

Solana ETF news today shows Goldman Sachs disclosed a $108M position while total SOL ETF inflows reached $1.45B. Analysts now expect up to $6B in institutional demand as Solana trades 71% below its all-time high.

Bitcoin ETF News Today: $2.1B Inflows Signal Strong Institutional Demand for BTC

Bitcoin ETFs news recorded $2.1B inflows over 8 consecutive days, marking one of the strongest recent accumulation streaks. Here’s what the latest Bitcoin ETF news means for BTC price and whether the $80K breakout level is next.

Michael Saylor: Winter is Over – Is He Right? 5 Key Data Points (2026)

Michael Saylor tweeted yesterday “Winter‘s Over.” It is short. It is bold. And it has the crypto world talking.

But is he right? Or is this just another CEO pumping his bags?

Let us look at the data. Let us be neutral. Let us see if the ice has really melted.

WEEX Bubbles App Now Live Visualizes the Crypto Market at a Glance

WEEX Bubbles is a standalone app designed to help users quickly understand complex crypto market movements through an intuitive bubble visualization.

Polygon co-founder Sandeep: Writing after the chain bridge chain explosion

In three weeks, Drift, Hyperbridge, and KelpDAO were consecutively hacked, resulting in nearly $900 million in losses. Polygon's CEO wrote that the problem lies not with any single team, but with the "notary" style architecture shared by the entire industry—relying on one or two signers to stamp cro...

Major Upgrade on Web: 10+ Advanced Chart Styles for Deeper Market Insights

To deliver more powerful and professional analysis tools, WEEX has rolled out a major upgrade to its web trading charts—now supporting up to 14 advanced chart styles.

Morning Report | Aethir secures a $260 million enterprise contract with Axe Compute; New Fire Technology acquires Avenir Group's trading team; Polymarket's trading volume surpassed by Kalshi

Overview of Important Market Events on April 23

Why a Million-Follower Crypto KOL Chooses WEEX VIP?

Discover why top crypto KOL Carl Moon partnered with WEEX. Explore the WEEX VIP ecosystem, 1,000 BTC protection fund, and exclusive rewards for serious traders.

CoinEx Founder: The Crypto Endgame in My Eyes

The industry will not disappear, but it will shrink significantly.

Spark Coin (SPK): Explodes 73% as Aave Bleeds $15B, A Good Investment Now?

Spark coin (SPK) surged 73% as $15 billion fled Aave after the KelpDAO hack. This article explains what Spark is, why it’s pumping, and whether it is a good investment right now.