Nvidia pushes to build Nemotron 4 open-source model as it seeks to drive chip demand and compete with customers
Nvidia is stepping up its push into open-source AI by investing resources in an internally developed flagship model. The company hopes the model will drive demand for its hardware, but the effort could also put it in competition with its own customers and partners. Several people involved in the Nemotron project said Nvidia plans to build the largest foundation model in the Nemotron 4 series, with performance aimed at the world’s leading open-source models. A paper on Nvidia’s previous flagship model listed 570 authors, and employees said the team working on Nemotron 4 will be even larger. “At this stage, everyone wants to be involved,” a former employee said. Nvidia has recently released several open-source models, and the new effort builds on that work. On Tuesday, the company released Nemotron 3.5 Lightning, a lightweight model designed to run AI agents efficiently and quickly. Nvidia also introduced free model-routing software to help companies quickly build model-orchestration tools that assign different AI tasks to models suited to the work and offering the best cost efficiency. Several Nemotron employees said the flagship Nemotron 4 will have at least a trillion parameters. Parameters are the adjustable units that a model continually modifies during learning. That would make the model roughly twice the size of Nvidia’s current flagship, Nemotron 3 Ultra, which was released in June. Even at a trillion parameters, Nemotron 4 would still be smaller than leading U.S. open-source models. Nvidia nevertheless places heavy emphasis on model-compression technology, which could allow a smaller model to deliver better performance. Nvidia’s increased investment in computing capacity for training its own models highlights its commitment to the open-source strategy. The company leases AI servers from cloud providers that purchase Nvidia chips. As of April, the total value of Nvidia’s long-term cloud-services procurement agreements had risen to $28 billion, with contracts running through early 2031—about three times the amount disclosed a year earlier. Some of that computing capacity is also used for other research projects. The aggressive Nemotron effort puts Nvidia in a delicate position: It is effectively competing at the same time with several open-source startups in which it has invested and with major AI labs that are among its most important customers. OpenAI has consistently been one of the largest sources of demand for Nvidia chips, and Nvidia has invested $30 billion in the AI lab. Although the flagship Nemotron 4 is expected to fall short of the performance of closed, frontier models from OpenAI and Anthropic, open-source models are being adopted by more companies because of their cost advantages and are taking on some business use cases. Nvidia has also made large investments in several U.S. open-source companies working to promote their own large language models, including Reflection AI and Thinking Machines Lab. More than 10 current and former employees, Nvidia partners and Nemotron users said Nvidia’s core rationale for expanding in open source is that a richer open-source ecosystem can steadily increase demand for GPUs. Nvidia’s chip demand is currently concentrated among a small number of frontier labs and cloud providers, including OpenAI and SpaceX; some of those companies have already begun developing their own AI chips. Nvidia’s strategy is that it will benefit if large numbers of startups and traditional enterprises can use cost-effective open-source models optimized for Nvidia hardware to develop AI applications, or if the broader open-source sector accelerates innovation. “Whichever company builds a great open-source model, Nvidia is a winner,” said Anastasios Angelopoulos, chief executive of AI model-evaluation organization Arena. Nvidia is intensifying competition in open source An Nvidia executive said the company believes an internally developed large model can stimulate competition across the industry, encourage the development of more open-source models and ultimately increase GPU demand. Nvidia has created the “Nemotron Alliance,” bringing together several U.S. open-source developers to jointly provide training data and research ideas and to develop next-generation models. Kari Briski, Nvidia’s vice president of generative AI, said in an email: “Nvidia continues to invest in the Nemotron project because we believe countries and companies of all kinds need accessible, frontier open-source models to strengthen security, accelerate innovation and build a technology foundation that can be continually improved and relied upon over the long term.” The launch date for Nemotron 4 has not been determined. Several employees said Nvidia has settled on the flagship model’s pretraining data and basic architecture, but its precise parameter specifications and release date have not been finalized, and the final full-scale training process has not begun. Training could take several months. Two employees said the model could be completed in late fall this year, while others expected it to take longer. Brian Catanzaro, Nvidia’s vice president of applied deep-learning research, said on a podcast in January that the Nemotron project “is about the long-term future of the company and is critically important.” Nvidia’s large models have already been deployed by some companies. One example is Palantir, which announced in June that it was working with Nvidia to deploy Nemotron-series models for U.S. government customers. Nvidia’s models, however, still have not reached the industry’s top tier in overall performance. As of now, Nvidia’s strongest model, Nemotron 3 Ultra, ranked second among U.S. open-source models in agent capabilities and overall intelligence in multiple evaluations by Arena and Artificial Analysis, behind only Thinking Machines’ newly released Inkling. It still trails the top U.S. open-source models and ranks outside the top 40 globally. Nvidia’s investment in large-model research is constrained by the company’s overall capital-allocation limits. Its cloud-services budget includes $7 billion for the fiscal year ending in January 2028. Although that is far below the spending of OpenAI and Anthropic, it exceeds the total funding raised since its founding by most leading open-source labs. One employee said Nvidia is seeking to allocate more computing capacity to Nemotron. Members of the Nemotron Alliance include Reflection, Cursor, Thinking Machines and Mistral. These companies are advancing their own open-source projects while providing training data, evaluation support and model-design ideas for Nemotron 4. People familiar with the collaboration said startup Prime Intellect is providing 300,000 simulated environments for model training. AI coding startup Cognition is also an alliance member, although the partnership has not been announced publicly, and the two sides are discussing Cognition’s provision of code-training data to Nvidia. The partners have different reasons for joining the alliance. For companies developing their own large models, Nvidia’s open-source push can raise attention on the sector and expand the market. Others want to help build a shared model and eventually use the results directly, avoiding the high cost of training compute. Prime Intellect Chief Executive Vincent Weisser said the alliance is intended to combine the industry’s strength and prevent the market from being monopolized by a single “ultimate large model,” rather than having each company compete to build the strongest open-source model.
Several people involved in the Nemotron project said Nvidia plans to build the largest foundation model in the Nemotron 4 series, with performance aimed at the world’s leading open-source models. A paper on Nvidia’s previous flagship model listed 570 authors, and employees said the team working on Nemotron 4 will be even larger. “At this stage, everyone wants to be involved,” a former employee said.
Nvidia has recently released several open-source models, and the new effort builds on that work. On Tuesday, the company released Nemotron 3.5 Lightning, a lightweight model designed to run AI agents efficiently and quickly. Nvidia also introduced free model-routing software to help companies quickly build model-orchestration tools that assign different AI tasks to models suited to the work and offering the best cost efficiency.
Several Nemotron employees said the flagship Nemotron 4 will have at least a trillion parameters. Parameters are the adjustable units that a model continually modifies during learning. That would make the model roughly twice the size of Nvidia’s current flagship, Nemotron 3 Ultra, which was released in June. Even at a trillion parameters, Nemotron 4 would still be smaller than leading U.S. open-source models. Nvidia nevertheless places heavy emphasis on model-compression technology, which could allow a smaller model to deliver better performance.
Nvidia’s increased investment in computing capacity for training its own models highlights its commitment to the open-source strategy. The company leases AI servers from cloud providers that purchase Nvidia chips. As of April, the total value of Nvidia’s long-term cloud-services procurement agreements had risen to $28 billion, with contracts running through early 2031—about three times the amount disclosed a year earlier. Some of that computing capacity is also used for other research projects.
The aggressive Nemotron effort puts Nvidia in a delicate position: It is effectively competing at the same time with several open-source startups in which it has invested and with major AI labs that are among its most important customers. OpenAI has consistently been one of the largest sources of demand for Nvidia chips, and Nvidia has invested $30 billion in the AI lab.
Although the flagship Nemotron 4 is expected to fall short of the performance of closed, frontier models from OpenAI and Anthropic, open-source models are being adopted by more companies because of their cost advantages and are taking on some business use cases.
Nvidia has also made large investments in several U.S. open-source companies working to promote their own large language models, including Reflection AI and Thinking Machines Lab.
More than 10 current and former employees, Nvidia partners and Nemotron users said Nvidia’s core rationale for expanding in open source is that a richer open-source ecosystem can steadily increase demand for GPUs. Nvidia’s chip demand is currently concentrated among a small number of frontier labs and cloud providers, including OpenAI and SpaceX; some of those companies have already begun developing their own AI chips.
Nvidia’s strategy is that it will benefit if large numbers of startups and traditional enterprises can use cost-effective open-source models optimized for Nvidia hardware to develop AI applications, or if the broader open-source sector accelerates innovation.
“Whichever company builds a great open-source model, Nvidia is a winner,” said Anastasios Angelopoulos, chief executive of AI model-evaluation organization Arena.
Nvidia is intensifying competition in open source
An Nvidia executive said the company believes an internally developed large model can stimulate competition across the industry, encourage the development of more open-source models and ultimately increase GPU demand. Nvidia has created the “Nemotron Alliance,” bringing together several U.S. open-source developers to jointly provide training data and research ideas and to develop next-generation models.
Kari Briski, Nvidia’s vice president of generative AI, said in an email: “Nvidia continues to invest in the Nemotron project because we believe countries and companies of all kinds need accessible, frontier open-source models to strengthen security, accelerate innovation and build a technology foundation that can be continually improved and relied upon over the long term.”
The launch date for Nemotron 4 has not been determined. Several employees said Nvidia has settled on the flagship model’s pretraining data and basic architecture, but its precise parameter specifications and release date have not been finalized, and the final full-scale training process has not begun. Training could take several months. Two employees said the model could be completed in late fall this year, while others expected it to take longer.
Brian Catanzaro, Nvidia’s vice president of applied deep-learning research, said on a podcast in January that the Nemotron project “is about the long-term future of the company and is critically important.”
Nvidia’s large models have already been deployed by some companies. One example is Palantir, which announced in June that it was working with Nvidia to deploy Nemotron-series models for U.S. government customers.
Nvidia’s models, however, still have not reached the industry’s top tier in overall performance. As of now, Nvidia’s strongest model, Nemotron 3 Ultra, ranked second among U.S. open-source models in agent capabilities and overall intelligence in multiple evaluations by Arena and Artificial Analysis, behind only Thinking Machines’ newly released Inkling. It still trails the top U.S. open-source models and ranks outside the top 40 globally.
Nvidia’s investment in large-model research is constrained by the company’s overall capital-allocation limits. Its cloud-services budget includes $7 billion for the fiscal year ending in January 2028. Although that is far below the spending of OpenAI and Anthropic, it exceeds the total funding raised since its founding by most leading open-source labs. One employee said Nvidia is seeking to allocate more computing capacity to Nemotron.
Members of the Nemotron Alliance include Reflection, Cursor, Thinking Machines and Mistral. These companies are advancing their own open-source projects while providing training data, evaluation support and model-design ideas for Nemotron 4. People familiar with the collaboration said startup Prime Intellect is providing 300,000 simulated environments for model training. AI coding startup Cognition is also an alliance member, although the partnership has not been announced publicly, and the two sides are discussing Cognition’s provision of code-training data to Nvidia.
The partners have different reasons for joining the alliance. For companies developing their own large models, Nvidia’s open-source push can raise attention on the sector and expand the market. Others want to help build a shared model and eventually use the results directly, avoiding the high cost of training compute.
Prime Intellect Chief Executive Vincent Weisser said the alliance is intended to combine the industry’s strength and prevent the market from being monopolized by a single “ultimate large model,” rather than having each company compete to build the strongest open-source model.
