AI Hype vs Reality: What 150+ Mathematicians Want Policymakers to Know in 2026
Mathematicians call for policymakers to distinguish AI hype from reality, urging evidence-based regulation.
Mathematicians warn governments not to buy AI hype: what’s really going on
More than 150 mathematicians have signed a joint declaration urging governments to be sceptical about sweeping claims of AI capability and to consult independent scientists before making policy. The trigger was a high-profile claim that an AI system had independently disproved an Erdős conjecture – a claim the signatories view as overstated and hard to validate.
The Reddit discussion that sparked this piece is here: r/ArtificialInteligence thread. The summary article cited in that post is from Futurism: link.
What triggered the declaration: AI, Erdős, and extraordinary claims
The declaration responds to marketing around an AI system allegedly generating a novel mathematical result without human help. Mathematicians say such claims are easy to oversell and extremely hard to check.
“Distinguishing flawed AI arguments from correct proofs is extremely difficult.”
In mathematics, a proof is binary: either correct or not. Large language models (LLMs) – the transformer-based systems powering modern chatbots – are powerful pattern predictors, not theorem provers. They are prone to hallucinations (plausible but false outputs) and often require careful tooling, such as proof assistants (software like Lean or Coq that verify every logical step), to achieve verifiable results.
Why AI-generated “proofs” are hard to trust without formal verification
Today’s strongest AI systems can assist mathematicians – searching literature, generating lemmas, or exploring conjectures. But without a formal end-to-end check, their outputs are closer to a sketch than a proof.
- Verification gap: Natural-language arguments can hide errors that only formal proof checkers will catch.
- Reproducibility: If a result relies on proprietary models, prompts, or toolchains, independent replication is hard.
- Attribution: It is difficult to tell whether a “new” result was inferred from training data, guided by human hints, or genuinely discovered.
- Benchmark mismatch: Success on benchmark problem sets doesn’t guarantee reliability on open, unsolved problems.
None of this means AI cannot contribute to mathematics. It can – and increasingly does – as a collaborator. But claims of autonomous discovery must meet a higher bar: formalisation, open artefacts, and independent validation.
What UK policymakers and public buyers should take from this
The UK has been pushing for evidence-led AI policy, including the creation of the AI Safety Institute and sector regulators that can issue guidance. This declaration reinforces a simple point: buy capability, not claims.
Procurement and policy guardrails
- Require verifiable evidence: For scientific or safety-critical claims, ask for formal proofs or independent replications, not just whitepapers.
- Mandate transparency: Log prompts, code, datasets (or at least data summaries), and compute used. Without artefacts, audits are theatre.
- Insist on domain expert review: Engage external mathematicians/scientists to evaluate technical assertions, not just vendor demos.
- Risk-tiering: Higher-risk deployments (health, policing, defence, welfare) demand stronger testing, bias audits, and human oversight.
- Data protection: Ensure compliance with UK GDPR/Data Protection Act 2018 – lawful basis, data minimisation, and DPIAs for high-risk processing.
For policing and surveillance, the bar should be especially high. Live biometric systems and predictive analytics carry significant rights risks. In defence, “human in/on the loop” requirements and rigorous test ranges remain essential.
A practical checklist for evaluating AI “breakthrough” claims
- What exactly is the claim? State the task, inputs, and definition of success in one paragraph.
- Evidence type: Is there a formal proof, a machine-checkable certificate, or just a narrative explanation?
- Reproducibility: Are code, prompts, seeds, and model versions available to independent auditors?
- Baselines and ablations: How does it compare to strong non-AI or simpler baselines? What parts are essential?
- Independence: Has an unaffiliated lab replicated the result?
- Limitations: Where does it fail? On which distributions or constraints?
- Safety and bias tests: For deployed systems, show red-team results and post-deployment monitoring plans.
Developers and researchers: use AI as a co-pilot, not a magic oracle
For engineers and scientists, the pragmatic path is clear:
- Integrate with formal tools: Couple LLMs to proof assistants and symbolic solvers; demand machine-checkable outputs for critical steps.
- Prefer retrieval-augmented generation (RAG) for factual work: RAG grounds model answers in documents, reducing hallucinations when sources are authoritative.
- Use unit tests and property checks: Treat AI-generated code and maths like any untrusted contribution – test it.
- Version and log everything: Prompts, model versions, data snapshots – so you can reproduce and audit.
Jargon quick hits:
- Transformer: The neural network architecture behind most modern LLMs.
- Hallucination: A confident but incorrect model output presented as fact.
- RAG (retrieval-augmented generation): Technique that fetches relevant documents and feeds them to a model to ground its replies.
- Alignment: Methods to make model behaviour match human intentions and values.
The wider risks flagged: surveillance, defence, and research integrity
The declaration also calls for regulation in military and mass-surveillance contexts, and highlights worries about unauthorised use of academic works and a rise in fake scientific papers. For the UK:
- Surveillance: Deployments of facial recognition or predictive tools need clear legal bases, necessity and proportionality tests, and independent oversight.
- Defence: Maintain strict human control, robust testing, and incident reporting for autonomous and semi-autonomous systems.
- Research integrity: Journals and universities should update policies to detect AI-generated fabrications, require code/data availability where possible, and sanction misuse.
- Copyright and licensing: Training or generation that uses academic content should respect rights and licences; document data sources and permissions.
Keep perspective: AI is useful – within verified boundaries
It’s possible to be bullish about AI’s practical value while sceptical of marketing hype. In maths and science, the right frame is “AI as lab assistant”: generate ideas, search literature, draft outlines – then verify rigorously.
For businesses and public bodies, the productivity case is strongest where outputs can be checked against ground truth: coding with tests, summarising known documents, data cleaning, and workflow automation with audit trails.
Environmental claims deserve the same scrutiny
Another area where hype meets reality is environmental impact. Water and energy use of AI data centres is often misunderstood. For a grounded look at water-cycle claims and cooling, see my explainer: AI, data centres and water: what actually happens.
Bottom line for the UK
Mathematicians are not saying “don’t use AI”. They’re saying: demand evidence that is checkable, reproducible, and independent – especially when claims are extraordinary. For policymakers, regulators, councils, the NHS, and police forces, that means tightening procurement standards, funding independent evaluations, and focusing on proven, auditable value over headline-grabbing demos.
If a vendor says their model discovered a new theorem, the right response is simple: show the formal proof, release the artefacts, and let independent experts verify it.
Related
Keep reading
AI
What Nolan's Anti-AI Film Tells Us About the Limits of AI Cost Cutting in Creative Industries
A film industry debate around Christopher Nolan, AI cost cutting and human craft reveals a useful lesson for UK creative businesses: AI can reduce some production costs, but it is not a substitute for taste, trust or a .
JoshuaJuly 26, 2026
AI
Why Reddit's reported Google AI access rethink matters for publishers and AI training
A reported rethink over Google's access to Reddit content shows how valuable human-written data has become for AI training, search visibility and publisher strategy.
JoshuaJuly 26, 2026
AI
Why AI Data Centres Are Facing Backlash Over Water, Power and Planning
AI data centres are no longer just a technology story. They are becoming a planning, utilities and public trust issue, with lessons for UK councils, businesses and AI policy.
JoshuaJuly 19, 2026
Tagged
Last updated
Category
aiLikes
Star Rating
No ratings yet
Comments
No comments yet - start the conversation.