Experts are surprised by how quickly artificial intelligence is tackling difficult mathematical problems, prompting some to wonder if human involvement will remain necessary.

Adrià Voltà

I am trying to resolve a mathematical puzzle that has challenged many of history’s top minds. I lack formal training in the subject beyond an old undergraduate physics qualification, making success unlikely. Yet I have an advantage: a computational tool capable of generating obscure insights rapidly. I submit a brief query on a specialized number theory conjecture and wait.

The tool is GPT 5.5 Pro, OpenAI’s newest model. For mathematicians, current AI systems show remarkable capability. Even amid fast development, gains in mathematical performance have been striking. Within months, leading figures have shifted from doubt to bold forecasts, privately discussing career impacts and whether certain projects are still worthwhile if machines might solve them sooner.

In April, I attended a quickly arranged gathering in San Francisco between mathematicians and AI specialists. The atmosphere combined enthusiasm with unease. If an untrained person could generate proofs instantly, what would that imply for specialists? Would human mathematicians still be needed? Could machines solve questions beyond human reach? The outcomes could reshape a practice dating back thousands of years, leaving little time for preparation.

Jacob Tsimerman of the University of Toronto, who helped arrange the event, stated that AI will arrive substantially and transform the discipline.

Views differ. Jeremy Avigad of Carnegie Mellon University recently wrote that few areas remain protected and that AI will soon outperform humans at proving theorems.

Some welcome the automation. Terence Tao of the University of California, Los Angeles, has described a shift from scarce proofs to plentiful ones, allowing once-intractable issues to be resolved by AI. Focus may move from discovering proofs first to understanding them first.

Mathematicians have long explored AI, yet only recently has it delivered practical results. Early efforts used custom neural networks for specific tasks. These systems proved hard to generalize and attracted limited interest.

Even after ChatGPT appeared in 2022, many remained skeptical, as early models handled basic arithmetic poorly and produced incorrect answers on advanced problems. As models grew larger and incorporated more mathematical training data, performance improved.

A key indicator came when systems attempted the International Mathematical Olympiad, a contest with six highly challenging questions. Researchers viewed strong results as a benchmark, expecting it would take years to achieve. In July 2024, Google DeepMind’s AlphaProof solved four questions, matching silver-medal level. A year later, both Google and OpenAI reached gold-medal performance using broader models. The outcomes prompted reevaluation. Ravi Vakil of Stanford University noted that perspectives changed markedly.

These abilities soon reached public tools and extended past competitions into research areas.

Credit:
https://www.newscientist.com/article/2526650-a-golden-age-of-maths-is-dawning-and-mathematicians-are-freaking-out/?utm_campaign=RSS%7CNSNS&utm_source=NSNS&utm_medium=RSS&utm_content=home
BCN