The contest between open-weight and closed AI models is changing faster than many expected. Based on recent releases, hardware investments, and usage data, I believe open models are becoming a serious commercial force.
Closed systems still lead in several areas. Yet their grip is weakening as open models get smarter, cheaper, and easier to run locally. That shift matters because it could change who controls AI, who pays for it, and where private data goes.
Hardware Strategy Reveals the Direction
OpenAI recently demonstrated early results from Jalapeno, its reported inference chip. The tests showed performance gains of up to 104 times on public models. Public testing matters because other companies can check those claims.
Inference is the process that produces an answer after a user enters a prompt. Jalapeno does not appear designed to train new models. OpenAI may still depend on Nvidia for training hardware in the near term.
Even so, the message is clear. Major AI companies want greater control over their computing systems. Meta is developing chips, while Google already uses its own tensor processing units.
This trend could reduce Nvidia’s inference business with the largest AI labs. Nvidia, however, may have a smart response.
Nvidia May Be Betting on Open Models
The Information reported that Nvidia agreed to buy Hugging Face. Neither company had publicly confirmed the deal at the time of the discussion, so caution is required.
Hugging Face acts much like GitHub for AI. Developers publish open-model weights there, while users can rent computing resources to run them.
If the reported purchase is real, I see three strategic benefits for Nvidia:
- It would gain direct access to a large open-model developer community.
- It could sell inference services without relying as heavily on major closed-model companies.
- It would profit as more businesses choose open models hosted in the cloud.
This would not be a simple software acquisition. It would be a bet that open-weight AI will capture a larger share of real business use.
The Usage Data Needs Careful Reading
Vercel CEO Guillermo Rauch shared data showing a sharp change in token use through the company’s AI gateway. Two months earlier, closed models handled 71.6% of tokens. Open models accounted for 28.4%.
More recent figures put open models at 62% of token volume, compared with 38% for closed systems. That sounds decisive, but requests tell another story.
Closed models still received 62% of actual requests. Open models received 38%. The token lead may partly reflect open systems using more tokens per request.
“We are seeing more and more requests recently going to more of the open-weight models.”
That distinction weakens claims of a complete market reversal. It does not erase the trend. Open-model demand is rising, even if the headline token numbers overstate its current reach.
Cost and Local Control Could Decide the Race
New releases show why adoption is growing. GLM 5.3 Flash reportedly scored 63.4 on the DeepSuite coding benchmark, beating Claude Opus 4.8. Its combined intelligence score placed it near Qwen 3.8 Max, but at a lower operating cost.
Qwen 3.8 Flash also offers strong performance, though its 125 billion parameters make cloud use more practical for most people.
Local hardware is reducing that barrier. Apple’s announced M5 Ultra configurations offer up to 512 GB of unified memory. Such machines could run larger models without sending every request to an outside provider.
The price remains steep. A 256 GB configuration was described as costing almost $11,000. Local AI is therefore not yet equally accessible. Businesses, researchers, and wealthy enthusiasts will benefit first.
Still, the direction should concern closed-model providers. Privacy, predictable costs, and direct control are powerful reasons to run AI locally.
Open Does Not Mean Closed Is Finished
Closed models retain major advantages in ease of use, support, safety controls, and access to top-tier performance. Many customers prefer a managed service over maintaining models and hardware.
Yet the old assumption that open models would remain years behind no longer looks safe. As the observer noted, the gap has moved from distant to “this close” within only a few years.
Businesses should test open and closed options against their own tasks, costs, and privacy needs. They should also avoid choosing providers from benchmark headlines alone. The next AI winner may not own the smartest model. It may own the cheapest, most trusted path for running many models.
Frequently Asked Questions
Q: What is an open-weight AI model?
It is a model whose trained numerical weights are available for others to download, host, inspect, or adapt under its license.
Q: How does an open-weight model differ from open-source software?
Open weights do not always include training data, full source code, or unrestricted rights. Users should review each model’s license.
Q: Why would Nvidia want Hugging Face?
The platform could connect Nvidia directly with developers and organizations that need cloud computing for open models. The reported deal remains unconfirmed.
Q: Are open models already more popular?
They led recent Vercel token volume, but closed models still led the number of requests. Adoption is growing, though market leadership is not settled.
Q: Can most people run large models at home?
Smaller models can run on consumer devices. Larger systems need costly memory and processing hardware, making cloud hosting more practical for many users.






















