Nvidia weighs aggressive plan to cut Rubin Ultra’s memory capacity
Nvidia is considering an aggressive plan to equip its next-generation Rubin Ultra GPU with less memory than originally planned, as supplies of advanced high-bandwidth memory chips remain tight. Three people directly involved in testing said Nvidia has tested at least three versions of the Rubin Ultra over the past several weeks. Some models had less memory than the specifications originally disclosed publicly. One reason Nvidia is considering a lower-memory version is that it may be unable to procure enough high-end memory to support the original design. Reducing memory capacity could affect chip performance, although Nvidia may be able to offset some of the shortfall through other technologies. AI companies running large models on a lower-memory version would generally need to deploy more GPUs. The change reflects how the chip industry’s capacity has struggled to keep pace with the explosive growth in data-center demand during the AI boom. Memory chips used alongside GPUs remain in short supply, and prices continue to rise. Higher costs have triggered ripple effects across the technology industry, pushing companies’ hardware spending above budget and prompting other vendors to raise prices for end products. Nvidia declined to comment. The situation is ironic: Strong demand for Nvidia GPUs in recent years has generated enormous demand for memory, contributing to supply shortages that are now forcing Nvidia itself to reconsider its memory configuration. The three Rubin Ultra samples currently in testing have lower memory specifications than the Rubin chips already in mass production and delivered to customers. Two Nvidia customers said the memory reduction may not significantly weaken demand for the Rubin Ultra. One said a lower-memory chip would probably be priced below the original design, offering a more cost-effective option that would help companies control expenses. The other customer said extremely large memory capacity is essential for running the most advanced frontier models, but that the company is less focused on the memory specification of an individual chip than on its long-term relationship with Nvidia across multiple generations of hardware products. Nvidia could also offset some of the disadvantages of reduced memory through its server-system design. Faster interconnect networks and more efficient data-storage solutions would allow customers to split large models and computing workloads across multiple Rubin Ultra chips running in concert. The lower-memory test configurations contrast with Nvidia’s earlier optimistic public comments. In a media briefing in mid-July, Andrew Bell, Nvidia’s senior vice president of hardware engineering, said the company had a dedicated team that anticipated and addressed supply-chain risks years in advance. “We planned ahead to address memory supply issues, and we will not be constrained by capacity in the short term,” Bell said. “Of course, the entire world is facing pricing pressure, and rising costs may be the bigger challenge. But on the supply side, we already have safeguards in place.” There is still time to make changes. Three people familiar with the matter and one Nvidia customer said the Rubin Ultra’s full hardware specifications have not been finalized. The chip is not expected to begin shipping until the end of next year, leaving Nvidia ample time to adjust the design based on memory availability, costs and customer demand. Industry publication SemiAnalysis first disclosed to customers last week that the memory specifications could be reduced. Pricing for the Rubin Ultra has not been determined. Data from Epoch AI shows that high-bandwidth memory accounts for more than half of the component cost of high-end AI chips. A lower-memory version could be priced substantially lower while remaining suitable for most AI workloads. Nvidia CEO Jensen Huang first unveiled the Rubin Ultra at the company’s 2025 annual developer conference. He said each GPU would have 1TB of HBM4E memory, a new generation of high-bandwidth memory for AI chips that offers improvements in data throughput, storage capacity and energy efficiency over the previous-generation HBM4. Slides shown at the event said the 1TB of memory would consist of 16 memory stacks. People familiar with the testing said Nvidia’s current versions reduce the number of memory-stack layers and the memory capacity of each die. Some samples even use older HBM4. The lowest-capacity test sample has just 192GB of total memory, while another version has 256GB—well below the originally planned 1TB. By comparison, Nvidia’s current Vera Rubin chip supports up to 288GB of HBM4 memory. Industry participants said lowering the memory capacity of each card would allow Nvidia to allocate limited memory resources across more GPUs, helping stabilize overall chip production. Sources said a major reason Nvidia is considering the downgrade is that memory makers may be unable to produce enough HBM4E by the Rubin Ultra’s mass-production schedule. HBM4E requires higher-density memory cells and high-speed circuit interconnects, while integrating it into chip packages is also significantly more difficult. Nvidia has taken steps to ease pressure on memory supplies. Late last month, Nvidia and the parent company of SK hynix reached a $500 billion agreement to jointly develop multiple versions of the next-generation HBM4 and HBM4E memory technologies. Raj Milpuri, Nvidia’s vice president of global AI cloud and infrastructure, said at a media briefing in July that the agreement was intended to “ensure a stable supply of high-bandwidth memory.” He added that SK hynix would invest in expanding high-bandwidth-memory capacity, with production prioritized for Nvidia. SK hynix announced in June that it planned to double its memory-chip capacity over the next five years.
Three people directly involved in testing said Nvidia has tested at least three versions of the Rubin Ultra over the past several weeks. Some models had less memory than the specifications originally disclosed publicly. One reason Nvidia is considering a lower-memory version is that it may be unable to procure enough high-end memory to support the original design.
Reducing memory capacity could affect chip performance, although Nvidia may be able to offset some of the shortfall through other technologies. AI companies running large models on a lower-memory version would generally need to deploy more GPUs.
The change reflects how the chip industry’s capacity has struggled to keep pace with the explosive growth in data-center demand during the AI boom. Memory chips used alongside GPUs remain in short supply, and prices continue to rise. Higher costs have triggered ripple effects across the technology industry, pushing companies’ hardware spending above budget and prompting other vendors to raise prices for end products.
Nvidia declined to comment.
The situation is ironic: Strong demand for Nvidia GPUs in recent years has generated enormous demand for memory, contributing to supply shortages that are now forcing Nvidia itself to reconsider its memory configuration. The three Rubin Ultra samples currently in testing have lower memory specifications than the Rubin chips already in mass production and delivered to customers.
Two Nvidia customers said the memory reduction may not significantly weaken demand for the Rubin Ultra. One said a lower-memory chip would probably be priced below the original design, offering a more cost-effective option that would help companies control expenses.
The other customer said extremely large memory capacity is essential for running the most advanced frontier models, but that the company is less focused on the memory specification of an individual chip than on its long-term relationship with Nvidia across multiple generations of hardware products.
Nvidia could also offset some of the disadvantages of reduced memory through its server-system design. Faster interconnect networks and more efficient data-storage solutions would allow customers to split large models and computing workloads across multiple Rubin Ultra chips running in concert.
The lower-memory test configurations contrast with Nvidia’s earlier optimistic public comments. In a media briefing in mid-July, Andrew Bell, Nvidia’s senior vice president of hardware engineering, said the company had a dedicated team that anticipated and addressed supply-chain risks years in advance.
“We planned ahead to address memory supply issues, and we will not be constrained by capacity in the short term,” Bell said. “Of course, the entire world is facing pricing pressure, and rising costs may be the bigger challenge. But on the supply side, we already have safeguards in place.”
There is still time to make changes.
Three people familiar with the matter and one Nvidia customer said the Rubin Ultra’s full hardware specifications have not been finalized. The chip is not expected to begin shipping until the end of next year, leaving Nvidia ample time to adjust the design based on memory availability, costs and customer demand. Industry publication SemiAnalysis first disclosed to customers last week that the memory specifications could be reduced.
Pricing for the Rubin Ultra has not been determined. Data from Epoch AI shows that high-bandwidth memory accounts for more than half of the component cost of high-end AI chips. A lower-memory version could be priced substantially lower while remaining suitable for most AI workloads.
Nvidia CEO Jensen Huang first unveiled the Rubin Ultra at the company’s 2025 annual developer conference. He said each GPU would have 1TB of HBM4E memory, a new generation of high-bandwidth memory for AI chips that offers improvements in data throughput, storage capacity and energy efficiency over the previous-generation HBM4. Slides shown at the event said the 1TB of memory would consist of 16 memory stacks.
People familiar with the testing said Nvidia’s current versions reduce the number of memory-stack layers and the memory capacity of each die. Some samples even use older HBM4. The lowest-capacity test sample has just 192GB of total memory, while another version has 256GB—well below the originally planned 1TB. By comparison, Nvidia’s current Vera Rubin chip supports up to 288GB of HBM4 memory.
Industry participants said lowering the memory capacity of each card would allow Nvidia to allocate limited memory resources across more GPUs, helping stabilize overall chip production.
Sources said a major reason Nvidia is considering the downgrade is that memory makers may be unable to produce enough HBM4E by the Rubin Ultra’s mass-production schedule. HBM4E requires higher-density memory cells and high-speed circuit interconnects, while integrating it into chip packages is also significantly more difficult.
Nvidia has taken steps to ease pressure on memory supplies. Late last month, Nvidia and the parent company of SK hynix reached a $500 billion agreement to jointly develop multiple versions of the next-generation HBM4 and HBM4E memory technologies.
Raj Milpuri, Nvidia’s vice president of global AI cloud and infrastructure, said at a media briefing in July that the agreement was intended to “ensure a stable supply of high-bandwidth memory.” He added that SK hynix would invest in expanding high-bandwidth-memory capacity, with production prioritized for Nvidia. SK hynix announced in June that it planned to double its memory-chip capacity over the next five years.
