Kids outlearn AI—and we still don’t know why

2026-08-26 · MIT Technology Review

Kids Outlearn AI: The Unsolved "Data Efficiency Gap"

For at least 100,000 years, human children were the only entities capable of learning a language to perfect fluency. Today, while Large Language Models (LLMs) like Claude, DeepSeek, and GPT models can converse naturally, a stark contrast remains behind the computational curtain: the massive amount of data required to train them.

The Data Efficiency Gap

This vast difference in required learning material between children and machines is known as the "data efficiency gap."

  • The Scale of AI Data: To achieve fluency, modern LLMs consume hundreds of thousands of times more words than a human child does. For instance, Meta's Llama 3.1 was trained on 15 trillion tokens. Frontier models may use ten times that amount. If printed on paper, the text used to train a modern LLM would stack past the International Space Station.
  • The Scale of Human Data: In contrast, a preteen raised in a linguistically rich environment hears only about 100 million words (or up to 300 million by age 20 with literacy). This amount of text would stack up to just 20 meters.
  • Early Milestones: Toddlers typically begin producing grammatically correct sentences after hearing between 10 million and 30 million words. Training an early model like GPT-2 on 30 million words merely produces a nonsense generator, not a fluent speaker.

As Stanford cognitive scientist Michael C. Frank puts it, while LLM progress is amazing, we currently have to "burn down a forest and scrape the entire sum of all human knowledge" to recreate a milestone that children achieve in their living rooms within a single year.

Why Reverse-Engineering Kids' Learning Matters

Understanding how children learn so much from so little could revolutionize both AI and cognitive science:

1. Overcoming Data Scarcity: The internet's supply of easily available training data could run dry as early as the 2030s. More data-efficient models are essential.

2. Practical AI Applications: Highly efficient models could improve AI training on video and help create chatbots for minority language communities.

3. Solving Cognitive Science Mysteries: Testing human learning hypotheses in machines could settle long-standing debates about the human mind, such as whether language processing is a biological quirk or reflects universal constraints.

The Debate Over Language Acquisition

How infants infer the infinite depths of language from a mere drop of experience remains a mystery. This puzzle lies at the heart of a decades-old debate:

  • The Innate Instinct (Chomsky): In the 1950s, MIT linguist Noam Chomsky argued that children are born with innate, hardwired knowledge of grammar. He pointed to the "poverty of the stimulus," arguing that syntax is too complex and children's exposure too limited to be learned purely from statistical experience.
  • The Environmental View (Skinner): Psychologist B.F. Skinner countered that language acquisition is entirely environmental, learned through conditioning and reinforcement.

By testing these theories using machine models, scientists hope to finally understand the miraculous way children master language.

Source