Beta The Briev beta is out. Free on iPhone via TestFlight — install it in under a minute.

Join the beta ↗
Briev
Live
Technology

AI struggles to recreate Einstein’s relativity in historic data experiments

Tests that limit AI training to pre-1911 knowledge show language models cannot independently derive general relativity, exposing current reasoning limits.

The idea of an “Einstein test” was introduced at the India AI Summit, where Demis Hassabis suggested training a large language model on all knowledge existing before 1911 to assess whether it could independently formulate general relativity. Subsequent efforts by multiple groups—including vintage models trained on pre-1900 data, a 1930-cutoff model, and the Ranke-4B series—have produced occasional hints of insight, such as a vague description of light quanta, but have largely fallen short of true scientific breakthroughs.

Papers by Tom Zahavy and Sendhil Mullainathan highlight that current AI excels at finding statistical correlations rather than making the abductive leaps Einstein used. Experiments with synthetic orbital-mechanics data showed models inventing incorrect gravitational laws instead of uncovering Newton’s law. While AI has demonstrated notable advances in mathematics, these successes still rely on recombining existing ideas rather than originating novel concepts. Overall, the experiments illustrate that without redesigning core reasoning architectures, AI is unlikely to replicate the creative leaps that produced relativity.

Why it matters

Understanding AI’s reasoning limits helps gauge its potential for genuine scientific breakthroughs.

In this story

einstein testartificial general intelligencevintage language modelabductive reasoningorbital mechanics modelhistorical LLMAI limitationsscientific discovery
Get the beta ↗