These startups are chasing the next big thing in LLMs
2026-08-20 · MIT Technology Review
These startups are chasing the next big thing in LLMs
MIT Technology Review’s What’s Next series looks across industries, trends, and technologies to give readers a first look at the future.
The Foundation Begins to Age
In the summer of 2017, Google AI researchers published “Attention Is All You Need,” introducing the transformer neural network. It proved exceptionally capable at processing long sequences of data, particularly text.
Nine years later, transformers power every major large language model on the market. “The entire AI industry is built on transformers,” says Justin Dangel, cofounder and CEO of Subquadratic. “They are one of the most important innovations in the history of computer science, and they’ve changed the world.”
Yet transformers are beginning to show their limitations. Many recent advances in LLMs—such as reasoning models and the ability to handle massive inputs—are not clean extensions of the core technology but workarounds that patch fundamental flaws.
A growing group of scientists and engineers is now asking what comes next. LLMs are here to stay, but the way they are built is up for grabs. MIT Technology Review has dubbed this next generation LLMs+.
A wave of startups has emerged hoping to push the boundaries of this technology. While some will undoubtedly fail, they have everything to gain and far less to lose than today’s leading companies.
The Problem with Dense Attention
The key strength of transformers is dense attention, which encodes the meaning of text by comparing every token with every other token through multiplication.
This mechanism captures meaning with remarkable accuracy. However, as text length grows, the number of computations increases quadratically. Processing a 10,000-word document may require 50 million multiplications. This is the primary reason LLMs consume enormous amounts of power.
The costs are staggering. OpenAI plans to spend $50 billion on computing this year alone. The International Energy Agency forecasts that data center electricity consumption will double by 2030.
Furthermore, transformers struggle with the demands of modern LLMs. Because they process text word by word, they are not well-suited for maintaining very large context windows. Yet harder tasks require ingesting entire document libraries, massive codebases, or, in the case of agents, outputs from other models.
Reasoning models exacerbate the issue by generating and rereading their own “chain of thought” notes, further increasing the volume of data that must be tracked.
As LLMs grow more powerful, the transformer architecture has become a bottleneck. Its greatest strength has turned into a critical limitation.
01: Rethinking Attention
One direct approach to making LLMs faster and cheaper is to replace dense attention with sparse attention, which performs calculations on only selected pairs of tokens rather than all possible pairs.
Over the years, researchers have proposed many sparse attention mechanisms, but none matched dense attention’s ability to capture meaning—until now, according to some startups.
Subquadratic, based in Miami, claims to have developed the first sparse attention mechanism that rivals leading LLMs on tasks including search and coding. Its model, SubQ, dynamically determines which words matter for each piece of text it processes. The company reports thousands have joined its waitlist and plans to release the model widely soon.
Manifest AI, based in San Francisco, takes a more radical approach by replacing attention entirely with a mechanism it calls power retention. Rather than tracking everything in the context window, it stores only the most relevant information for the current task and maintains a rolling summary, dropping less relevant data as new information arrives.
While sparse attention still keeps a rough representation of everything seen, power retention actively summarizes and prunes context. The basic retention concept has existed for a decade, but Manifest AI claims it has refined the technique sufficiently to compete with transformer-based models for the first time.
The company says transformer models can be converted to power retention models with minimal retraining. To demonstrate this, it transformed the open-source coding model StarCoder into PowerCoder and released Brumby, which it claims rivals certain versions of Alibaba’s Qwen model.
Manifest AI wants its power retention technology to become the new standard for building the next generation of LLMs.