(86) 0755 2103 0230 Address: Shenzhen,Guangdong Province,China
Home News In the final week of July, NVIDIA made headlines three times, and none of them were good news.
In the final week of July, NVIDIA made headlines three times, and none of them were good news.

Have you noticed this pattern: every time you scroll across NVIDIA news, it’s either “prices surge again” or “another arrest made”?


In the final week of July, NVIDIA made headlines three times in a row.


● Enterprise side: The flagship AI server priced at $9.1 million per unit is getting its memory cut to reduce costs.

● Smuggling side: A senior NVIDIA manager was found to have turned a blind eye in a chip smuggling case, with seven suspects detained.

● Consumer side: The RTX 5070 Ti has jumped 48% in half a year, with this mid-range graphics card now selling for 10,700 RMB.


On the surface, these three incidents seem unrelated. But connect the dots, and a shared truth emerges:


The frenzied price hikes of AI hardware have spread from cloud data centers all the way to consumer desktops, and profit distribution across the entire industrial chain is being forcibly reshuffled.


What’s more alarming than price hikes: miners earn less than those selling shovels


$9.1 Million Servers Are Getting Their Memory Halved

Let’s start with the most eye-opening story.


NVIDIA Vera Rubin NVL72 stands among the world’s highest-performance AI computing platforms. Unveiled this January, its launch specifications stunned everyone. When running Mixture-of-Experts models, Blackwell GPU systems deliver 80,000 Tokens per second under a 150-megawatt power draw, while the VR200 NVL72 multiplies this figure tenfold, hitting 800,000 Tokens per second.


Ten times faster. Not 10%, ten times.



But the optimism faded quickly. A Bernstein report crunched the numbers:A single NVL72 Rubin rack could cost as high as $9.1 million, up $1.3 million from the prior estimate of $7.8 million. The cause? Memory pricing.


When building the most advanced AI servers,the biggest cost component is not GPUs, but memory.


HBM4 memory prices may surge to $53 per GB by 2027.LPDDR5X memory for one VR200 unit can cost $1.2 million, accounting for 29% of its total bill of materials. To put that in perspective, 20% is normally considered an extreme ceiling.


A recent GF Securities report revealed:NVIDIA is drastically adjusting memory configurations for Vera Rubin systems. SOCAMM capacity is slashed from 192GB to 96GB — cut in half. Supporting memory for Vera CPUs drops from 54–55TB to 28TB, also reduced by half.


The consequence of this move:Trading performance for cost, instead of spending more to boost performance.


When even flagship AI servers must economize on memory, you can gauge how severe supply constraints have become across the whole industry.


Tightened supply chains bring not only price hikes, but also plenty of shadowy activity.


NVIDIA Manager Allegedly Cut Corners; Seven Suspects Detained

On July 28, the Keelung District Prosecutors Office released updates on the Supermicro AI server smuggling case, with additional suspects placed under detention.


Including previously detained persons, seven suspects are now held in connection with the case. Those arrested include a Supermicro sales manager, General Manager of Qunyun Technology, a senior executive at Chief Telecom, and one particularly notable figure — a senior NVIDIA manager.


Investigations found he deliberately overlooked compliance rules during customer eligibility reviews, issuing approvals to let restricted hardware pass screening procedures.


△ 50 seized high-end Supermicro AI servers waiting to be smuggled overseas


The scheme operated with a straightforward workflow: Orders placed via shell companies → Data center rental at IDCs to fabricate the illusion of local usage → Bypassing NVIDIA and distributor compliance audits → Customs declaration for export to Japan → Transshipment to restricted regions.


Simply put, the tactic relies on legitimate ordering, fake on-site deployment, and covert resale.


During the first round of raids this May, authorities intercepted more than 50 AI servers equipped with NVIDIA B300 chips. The case involves NT$700 million, equivalent to roughly 146 million RMB.


Valuable chips are enough to push a senior manager at a top tech firm to take enormous risks. This reveals a harsh reality:When supply shortages reach a critical level, gray markets emerge spontaneously.


Ultimately, ordinary end consumers end up footing the bill for supply deficits.


A Mid-Tier Graphics Card Rises by 3,500 RMB Within Half a Year

Now let’s turn to the consumer market.


According to BenchLife, NVIDIA recently issued another pricing adjustment notice to downstream supply chains, marking the third major price hike in 2026.


Prices went up once in January, again in May, and once more in July. Worse still, price increases have spread from high-end GPUs to mainstream mid-range and entry-level models. Previously, only flagship cards like the RTX 5090 became pricier, and regular users could shrug it off. Not anymore. The RTX 5070 Ti, equipped with 16GB VRAM and positioned as a premium mid-range sweet-spot GPU, now carries a retail price of 10,700 RMB.


Just last week, its price stood at 9,299 RMB. Back in November 2025, its launch price was 7,200 RMB. Over half a year, the price jumped by 3,500 RMB, a total increase of 48.6%.


Even more striking: rumors peg the starting price of the RTX 5080 at 12,000 RMB, 44.6% higher than its official founder’s edition MSRP of 8,299 RMB.


VRAM cost hikes alone cannot fully explain this inflation;panic buying is a major contributing factor. Major e-commerce platforms already removed multiple mainstream NVIDIA graphics cards from listings last week, and some distributors locked inventories, withholding stock. They are not unwilling to sell — they fear selling early and missing out on future price gains.


Meanwhile, Samsung raised DRAM pricing by approximately 20%. Standard 2GB GDDR7 memory chips now sell for $20 (roughly 145 RMB), while 3GB variants carry price quotes of $60 to $70. Capacity only rises by 50%, yet procurement costs surge over threefold.


High-end cards are becoming prohibitively costly for enterprises, mid-range models are unaffordable for gamers, and what about entry-level options?


Believe it or not, there is news on the low-end segment too: AMD launched the RX 9050, offered in a 4GB VRAM variant. You read that correctly — it is 2026, and new graphics cards with only 4GB VRAM are hitting the market.


This GPU features 16 compute units and 1024 FP32 cores running at 2600MHz, with theoretical performance less than half of the RX 9060 XT. Frankly speaking, this might have qualified as a basic entry-level card three or four years ago, yet in 2026 it offers barely adequate performance at best.


What’s more alarming than price hikes: miners earn less than those selling shovels

To be clear, rising prices, rampant smuggling and stripped-down entry-level GPUs are merely surface symptoms. Where lies the root cause?


AI capacity mismatch.


What does this mean? Demand already outpaces available supply, and upstream price surges essentially signal insufficient capacity to satisfy all buyers. But when companies tally their financial results, a critical problem emerges:


The gold miners earn less than merchants selling shovels.


This metaphor fits perfectly here. Cloud providers and AI startups act as the miners, while NVIDIA and Samsung are the shovel vendors. Shovel sellers now command exorbitant profit margins, and most profits earned by miners from their hard work ultimately flow to suppliers. Under these conditions, how many miners will keep digging?


△ Large language model ranking chart


This is not merely intuitive; hard data backs it up. Take Anthropic: annualized revenue hits 74.3 billion with gross margins exceeding 60%, which looks impressive on paper. But as a model developer, it still counts as one of the miners. What share of its earnings must be paid to NVIDIA and Samsung? Just consider HBM4 climbing to $53 per GB, and Samsung’s 20% price hike — every link across the supply chain feels the shockwaves.


This leaves the market with two clear paths forward:


  1. Improve miners’ return profile: Ensure builders of AI applications and cloud services can generate sustainable profits and maintain investment momentum. Yet this path appears unworkable. Capital expenditure data for cloud service providers reveals most spending flows straight to NVIDIA, leaving insufficient returns to attract continued investor funding.

  2. Resolve industrial capacity mismatches: This is the second option, and developments are already underway. Memory cuts on Vera Rubin systems, the emergence of gray-market smuggling, consumer GPU price hikes — all are manifestations of this shift. The entire industrial chain is forced to rebalance supply and demand, restoring reasonable profit margins for every participant.



AI capability stays unchanged; what shifts is market dynamics.AI capabilities keep advancing rapidly, yet the market has transitioned from blind, all-in investment into an era of rigorous financial scrutiny.


Once businesses start calculating costs and returns, the severe imbalance becomes impossible to ignore.


Three Events, One Interconnected Supply Chain

Connect these three incidents with the logic outlined above:


Memory price hikes → Escalating server costs → NVIDIA cuts memory specifications to lower expenses → Supply shortfalls fuel gray-market smuggling → Cost pressures trickle down to consumer hardware → Graphics card prices surge nearly 50% within half a year → Entry-level market shrinks to 4GB VRAM solutions


Beneath this chain lies a deeper imbalance: Miners see weaker returns than shovel sellers, and profit allocation across the whole industrial chain is undergoing forced restructuring.


Data centers downgrade hardware to cut costs; NVIDIA pursues internal wrongdoers; gamers debate whether to bite the bullet and upgrade. Meanwhile, memory manufacturers including Samsung and SK Hynix emerge as the real winners of this price surge cycle. A straightforward 20% price increase from them sends shockwaves rippling from cloud infrastructure all the way to consumer desktops.


At the 2026 AI boom feast, ordinary gamers are left facing ever-higher bills.


For the broader industry, a critical question hangs in the air: If all miners stop digging for gold, how long can the shovel merchants stay in business?




3DSTOR is a long-established global IT component supplier maintaining stable long-term partnerships with top-tier brands including Intel, AMD, NVIDIA, Western Digital and MSI. Its product portfolio covers servers, motherboards, graphics cards, CPUs, hard drives and more. Focused on serving global B2B markets, 3DSTOR delivers diverse AI/ML/HPC solutions and provides one-stop procurement services for IT hardware clients with varied requirements.