OpenAI published a paper on August 1 claiming an unreleased model called Astra produced ten new results in mathematics and theoretical computer science. Total inference cost, at the lab’s own API rates, came to about $2,000. The list includes the first explicit construction of a non-sofic group, a question open since 1999, a disproof of Connes’s rigidity conjecture, an improvement on sphere-packing bounds that had stood since 1978, and resolutions of three Erdős problems.
The people who would know took it seriously. Thomas Bloom, who maintains erdosproblems.com, called it big news. Timothy Gowers, a Fields medalist, said he would recommend one of the proofs to the Annals of Mathematics without hesitation. Sébastien Bubeck, who runs math research at OpenAI, called the results beautiful, which is what you would expect him to say. Noam Brown pointed out there were no Millennium Prize Problems in the batch, which is not.
OpenAI did not ask anyone to take its word for it
Along with the manuscript and the model’s reasoning walkthroughs, OpenAI shipped machine-checkable Lean 4 certificates to GitHub. Anyone with the Lean compiler can verify every proof mechanically, on their own machine, without access to Astra, without a demo call, without trusting a single claim in the press release. The assertion and the means of falsifying it arrived on the same day.
I have spent fifteen years selling technology to people whose job is to not get fooled. Fastly’s buyers were engineers who ran their own load tests before they would take a meeting about pricing. Oracle’s buyers had procurement organizations whose entire function was structured distrust. Every AI deal since 2023 has hit the same wall: the buyer asks how they would know the output is correct, and the two answers on offer are both weak. Benchmarks are selected by the vendor. Pilots are staffed by the vendor’s engineers, which makes them a test of the vendor’s engineers.
A verification artifact the customer runs themselves is a third answer, and it is rare because it is genuinely hard to build. It is also the only one of the three that survives a hostile procurement review.
Most enterprise work has no Lean compiler
The limit here is obvious and I want to state it before anyone over-reads the result. Mathematics has formal verification. Your pipeline forecast does not. There is no compiler that will tell you an AI-drafted account plan is correct, and there never will be.
But some enterprise work does have a checkable oracle sitting right there, and those are the deals to lead with. Code that passes a test suite the customer wrote. Reconciliations that tie to the customer’s general ledger. Contract clauses graded against the customer’s own playbook. Support answers checked against a policy document the customer supplied. In each of those the buyer can grade your output without asking your permission or scheduling your solutions engineer. Sell into those categories first and let the fuzzier ones ride on the trust you earn there.
Two things can be true. This is a real mathematical result with artifacts a stranger can verify, and it is still a lab announcing a breakthrough by press release about a model no outsider can touch. The Leiden Declaration, endorsed by the International Mathematical Union, exists partly because mathematicians got tired of that pattern. Astra has no public interface and no announced release date. OpenAI is giving 100,000 academic researchers free access through 2027, which is a good-faith gesture and a distribution strategy at the same time.
What this costs you to copy
Pick your product’s central claim and ask what a hostile customer could run, on their own hardware, without your help, that would prove you wrong. If the answer is a slide, you have a positioning problem no amount of discounting will fix.
OpenAI spent $2,000 on the compute and an unknowable amount on the 249 pages and the Lean files that let strangers check the work. The second number is the one that bought the credibility.
