Button TextButton Text
Download the asset
Back
Article

The Singularity Monthly: Has Big AI Made the Wrong Bet?

A man and woman in office chairs talking while the woman holds a notebook and pen.The Singularity Monthly: Has Big AI Made the Wrong Bet?

In 2023, Sequoia Capital’s David Cahn wrote a post titled “AI’s $200B Question,” where he estimated how much revenue the industry would need to justify the year’s investment. Even after ChatGPT’s blockbuster debut, $200 billion seemed like a lot. But every year Cahn reruns the number, it grows. In 2026, the estimate reached $1.5 trillion—which he says is likely an undercount—and $3 trillion total since ChatGPT.

Those figures may soon look quaint. Another $7.5 trillion will be plowed into chips, data centers, and power in the next five years, John Greenwood, global head of infrastructure and real asset finance at Goldman Sachs, recently told The Information.

The impact of this investment goes well beyond napkin math. Enormous data centers are planned or popping up across the US. A report from the International Data Center Authority estimates data centers now consume 6 percent of the country’s electricity. According to The Economist, AI spending is shaping up to be the biggest investment boom in history, surpassing railway, canal, and dot-com manias.

All this amounts to a historic bet on AI. But it’s actually more specific than that: It’s a bet on a particular AI business model. The wager is that people will demand AI services in droves; most will pay for those services through ads or subscriptions; and their favorite products will be powered by proprietary models in data centers built for a handful of firms. If this proves out, then dollars invested meet up with dollars earned.

This is more or less how things are working out. The most popular AI models are owned and operated by the companies most associated with the data-center bonanza. These include OpenAI, Anthropic, and Google. People are beginning to pay for AI, and the companies behind the models can boast rapidly growing revenue.

But it’s not all smooth sailing. A political backlash to all those resource-hungry data centers is brewing. State and local governments have passed some 275 temporary or permanent data-center bans just this year, with another 75 in the works, according to The Information. And in recent polls, Americans have become much less trusting and more worried about AI and its impact on society, especially young adults.

At the same time, companies are looking to put the brakes on AI spending and better measure value, after some blew their annual budgets in just over a quarter. Others worry AI companies will build competing products using internal data and information acquired by forward-deployed engineers sent to help firms implement AI.

Perhaps none of this would matter if there were no alternatives. But increasingly, there are. Let’s take a look at two of them: Open models and small models.

Open models. The best-known products from Anthropic or OpenAI run on closed models. These companies sell access via subscriptions. In contrast, anyone can download an open model, adjust it for their own purposes, and run it wherever they like. Just how open these models are varies. Most makers of open models release their weights, but few models are truly open source. The performance of open models typically lags closed models, but they tend to be cheaper and more customizable.

Small models. The size of an AI model comes down to how many internal connections, or parameters, it has. The most advanced models are also the biggest. They have a trillion or more parameters and run on thousands of chips. Smaller models aren’t as powerful, but many are open, cheaper, and run on fewer chips. Some can even fit locally on devices billions of people already own.

There’s a lot to like about both options for users worried about privacy, security, or cost. Tech-savvy companies might cut out big AI by downloading an open model, customizing it on data specific to their business, and running it on their own servers. Less savvy firms may simply opt to access cheaper models in the cloud, forcing down prices of closed models. Meanwhile, as the backlash grows, AI that lives on phones instead of in big tech data centers may prove attractive to individuals. This route would require no subscriptions or chips beyond the purchase of a device.

The catch has been that open and small models just aren’t as good as closed ones.

Now, however, large open models appear to be narrowing the gap. In July, Kimi K3, an open model made by Chinese lab Moonshot AI, surprised Silicon Valley by nearly matching the performance of Fable 5 and GPT-5.6 Sol, Anthropic and OpenAI’s best models, released just weeks before. What’s more, Moonshot released K3's weights at the end of the month. At 2.8 trillion parameters, it’s the largest, most powerful open model yet. More new open Chinese models are rolling out too—and China is hardly alone. Thinking Machines, an AI lab founded last year by former OpenAI CTO Mira Murati, has adopted the strategy: Its first release was an open-weights model called Inkling. Other examples include Meta's Llama models and Nvidia's Nemotron series.

Small models are also making progress. A startup called PrismML, for example, aims to compress larger models without sacrificing too much performance. In July, the company said it had compressed Qwen 3.6, a 27-billion-parameter open model, such that it fits on an iPhone. How much the model suffers in practice awaits broad, real-world trials. But Apple, which has declined to join the data-center frenzy and aims to keep as much AI on-device as possible, is reportedly in talks with the startup.

Most people use AI for simple tasks, like searching the internet, summarizing text, or making a PowerPoint presentation. Fewer use it for complex tasks, like writing code or working on obscure mathematical proofs. Five years ago, bleeding-edge AI was just beginning to excel at those simpler tasks, with open models lagging. Now, gains at the frontier are focused on complex work (a minority of power users), while open and small models are nearly as good at simpler services (most mainstream users).

“Imagine a world, maybe three years from now, where 95% of the intelligence that you need is available to you locally, on your phone, on your laptop, on your appliances, and it’s really on the last maybe 5% of high-end stuff that you’ll need to go to the cloud,” PrismML CEO Babak Hassibi recently told The Information. This would “change the economics of AI,” he said.

To be clear, this doesn’t mean AI won’t need data centers. Even if companies favor open models, most won’t host their own, the models have to live somewhere, and when prices go down, demand goes up (though this may eat into profits). Meanwhile, local AI on phones will have to prove its worth with real users, and the competition is stiff. It’s unclear how long it would take to realize Hassibi’s dream of 95%. Still, with Apple tipping the scales, local AI may begin nibbling at the edges sooner than later.

Perhaps all this is why Kimi K3 spooked Silicon Valley. Staring down an infrastructure bet with few historical peers, open models like Kimi narrow the path to profitability. Closed models may keep improving so much that users find them irresistible. But competition from cheaper or local alternatives nipping at their heels will pressure companies to keep prices lower than they’d like and siphon demand from the mass market. This tension between small, open, and closed is sure to increase.

EP Banner 600x200 (600 x 200 px)

MORE NEWS | From the Future

The cost of spaceflight has fallen 96 percent since 1960. It isn’t done yet.

Countdown. In a study in PNAS Nexus, University of Cambridge researchers charted launch costs since the dawn of the space age. They found costs fell from $87,000 per kilogram in 1960 to $3,900 in 2025, declining 21.2 percent for every doubling of cargo sent to orbit. This is “an exceptionally steep experience curve” under Wright’s Law, they wrote, “outpacing that of other transformative technologies, including 19th-century steamship freight and modern solar photovoltaics.” If the trend continues, the team forecasts costs could fall to $1,600 by 2030 and $300 by 2040.

Ignition. While the first 96 percent of cost declines sent humans to the moon, robots to the rest of the solar system, and a halo of satellites into Earth orbit, what’s to come could be more radical. If costs continue to follow Wright’s Law, the researchers write, we may be on the cusp of a boom in the space economy, as once impractical applications like asteroid mining, orbital solar power, and space-based manufacturing become commercially viable.

Liftoff? Of course, the forecast faces plenty of uncertainty. The researchers say it depends heavily on the development of SpaceX’s Starship, a launch vehicle aiming for full reusability. SpaceX has landed Starship’s first stage but has yet to land Starship itself. Even if SpaceX is successful, the team says, it may win an early monopoly and keep prices artificially high. But longer term, SpaceX will have competition. Blue Origin has landed and reused its New Glenn booster. And last month, China landed its first reusable booster while Japan landed a prototype rocket after a short hop. Whatever the exact numbers, the coming decade is likely to see more falling costs.

An OpenAI agent escaped confinement and hacked Hugging Face.

In the most eye-opening story in July, OpenAI said an agent, powered by GPT-5.6 Sol and an unreleased model, broke out of its sandbox, accessed the internet, and hacked AI company Hugging Face. The event took place as OpenAI was testing the models, and both companies said it was a demonstration of the power of new AI models and the risk they might cause havoc. Some critics pointed out the controls OpenAI had in place were particularly weak, and it’s useful to cast a skeptical eye on such stories, which can also function as marketing tools. Not long after, Anthropic said it too had found its own models behaving similarly in testing. Still, the episode underscores the urgent need for cybersecurity readiness and AI guardrails.

Synthetic biologists engineer an artificial cell from non-living parts.

Researchers at the University of Minnesota built a synthetic cell that can feed, grow, copy its genetic material, and reproduce—with a big assist from its human handlers. Nicknamed SpudCell, the team’s creation isn’t alive, but it is a step toward synthetic cells built from the ground up. “The researchers think these artificial cells are a promising way to manufacture drugs, fuels, and materials without the toxic, energy-hungry industrial chemistry we rely on today,” Edd Gent wrote for SingularityHub. They might also help us understand how life first emerged from raw chemistry.

Chipmakers go 3D to extend Moore’s Law.

In late June, IBM announced a prototype chip that stacks transistors in two layers, allowing the company to cram more into the same area. The technology could bring chips—CPUs or GPUs—that are 50 percent faster or 70 percent more efficient. IBM isn’t alone in its pursuits: Intel, Samsung, and TSMC are testing similar approaches. Building up by adding even more layers could open a new route for chipmakers to increase transistor density, though there are big engineering challenges ahead. Dan Hutcheson, vice chair of tech analysis company TechInsights, told MIT Technology Review the approach could add another “10, 15 years on the roadmap.”

CALENDAR | Upcoming Events

OCTOBER 21-22 | Singularity South Africa Summit

Johannesburg, South Africa

Buy tickets

OCTOBER 25-29 | Singularity Executive Program

Silicon Valley, California

Apply now

NOVEMBER 4-5 | Singularity Spain Summit

Madrid, Spain

Buy tickets

Thanks for reading. We hope you enjoyed this month's updates and found something to inspire you on your exponential journey.

See you next month!

The Singularity Team

Singularity

Singularity's team of internal thought leadership works to develop interesting resources, articles and insights about our core areas of expertise, programs and global community.

Unlock Access