DeepMind has published a case study on using its Gemini AI to evaluate mathematical problems listed in the Erdős Problems database, with the research paper reporting five seemingly novel autonomous solutions. The study was posted on arXiv.
The research matters because it tests the capacity of a large language model to contribute to research-level mathematics rather than benchmark-style puzzles. The Erdős Problems database contains unsolved conjectures posed by the mathematician Paul Erdős.
Automatica reported in July 2025 that an advanced version of Gemini Deep Think achieved Gold-medal standard at the International Mathematics Olympiad (IMO). The DeepMind blog notes that subsequent work built on that result.
The arXiv paper, titled “Semi-Autonomous Mathematics Discovery with Gemini: A Case Study on the Erdős Problems,” describes using a hybrid methodology to evaluate 700 conjectures labeled 'Open' in Bloom's Erdős Problems database. AI-driven natural language verification narrowed the search space, followed by human expert evaluation to assess correctness and novelty. The researchers addressed 13 problems marked as open: five through “seemingly novel autonomous solutions,” and eight through identifying previous solutions in existing literature. The paper says the 'Open' status of those problems appeared to stem from obscurity rather than difficulty.
The study also identifies issues in applying AI to mathematical conjectures at scale, including the difficulty of literature identification and what it terms the risk of “subconscious plagiarism” by AI. An addendum to the paper notes a reclassification of one problem as “Independent Rediscovery,” bringing the count of autonomous solutions down to five.
A social post summarizing the research claimed that AI had resolved nine open Erdős problems. That number is not supported by the primary research document. The arXiv paper states that five problems received novel solutions and eight had prior solutions identified; the paper's abstract and main text do not assert nine new resolutions. The DeepMind blog post describes an “extensive semi-autonomous evaluation … including autonomous solutions to four open questions” listed in the database, also differing from the social post’s figure.
The research is a preprint and has not yet undergone peer review. The DeepMind blog says that Level 2 (“publishable quality”) works from the project have been submitted to reputable journals, and that no Level 3 (“Major Advance”) or Level 4 (“Landmark Breakthrough”) results are claimed. The paper does not include independent test results beyond the methodology described.
Prompts and model outputs are available on GitHub. The case study was conducted using the Gemini model configured for semi-autonomous mathematics discovery, with human expert verification of results.