The Looming Proof: What an Unverified AI Math Breakthrough Means for Crypto's Hidden Infrastructure
NeoPanda
The headline arrives weightless. An AI system has solved three open problems in mathematics, and the crypto press repeats it as if proof were a weather forecast. I read the claim twice, then a third time, looking for the load-bearing detail. There is none. No model name. No paper. No Lean proof. No named mathematician confirming the verification. Just the narrative, already circling the market like a predator.
Narratives are liquid; truth is solid. In 2017, I watched a similar liquidity sweep through an ICO market that had convinced itself that whitepapers were optional. Now the same structure appears in a different costume: an unverified mathematical claim presented as a turning point for the AI-crypto convergence. The crowd wants the story. I want the math.
Let us calibrate against the benchmark that matters. FrontierMath, designed by Epoch AI, was announced with a blunt message: the best language models could solve a single-digit percentage of its problems. The test set is composed of mathematics at the research frontier, problems that require more than memorized patterns, more than fluent symbol shuffling. If a model truly crossed from single-digit performance to solving open research questions, we would be discussing a discontinuity large enough to break the current scaling paradigm, and large enough to reshape crypto's most sacred assumption: that verification is cheap, and that trust is a scarce resource.
For the crypto ecosystem, this is not idle curiosity. Every smart contract platform, every zero-knowledge rollup, every optimistic-proving mechanism rests on a series of mathematical claims: a solution can be verified, an invariant can be proven, absence of fraud can be declared sound. Bitcoin's proof-of-work is a repeated invocation of computational hardness assumptions; the chain exists only because verification is cheap and forgery is expensive. Automating mathematics is not a separate research program. It is a direct investment in the foundation of every token ledger, and in the edge of every auditor's report.
The first thing I did after seeing this story was to search for the three problems. That search failed. There is no public record of the statements, no pre-print repository, no formal library with a submission date. The only 'proof' on offer is the claim that a proof exists. In mathematics, a claim without a demonstrandum is not a theorem; it is a rumor. In markets, a rumor with momentum is a trade. The tension between those two definitions is where this analysis begins.
What does 'solved' actually mean in the community that owns the word? Mathematicians have a precise hierarchy. A conjecture is a proposition that no one has yet proven or disproven. A counterexample kills a conjecture. A proof is a chain of inference from accepted axioms to the proposition, auditable by any competent practitioner. Fermat's Last Theorem remained 'solved' only after Andrew Wiles produced a proof so deep that the first draft contained a fatal gap, requiring partnership with Richard Taylor to close. The Four Color Theorem was accepted only after the community accepted a computer-assisted proof decades after the first attempts, because humans could not manually verify the case analysis. The lesson is structural: mathematical truth is not the same as a declared result. It is the outcome of a social and technical verification process, and that process has never been fast.
If an AI solved three genuinely open problems, the evidence should look a particular way. There should be either a human-readable proof, reviewed by specialists, or a formal proof object in a system like Lean, Coq, or Isabelle, mechanically checked by code that has itself been audited. There should be notes on the method, on the compute budget, on the model architecture. The absence of all of these elements tells me something important: this is not yet a scientific event. It is a narrative event wearing a lab coat.
This is where I want to be careful, because I have a personal history with premature claims. In late 2017, at age twenty-five, I spent weeks modeling the token economics of a prominent ICO against its stated computational utility claims. I found a flaw in the reward distribution mechanism that ignored transaction fee volatility. The market did not care. The token surged anyway. My correction later proved out, but the lesson was not about being right; it was about knowing that the market's reward function does not include truth. Math does not care about your conviction. The market cares even less. The only antidote is to separate the claim from the evidence, and to hold the claim to the standard of the evidence.
Let us now ask the question that matters: if the event is real, what is the most plausible technical route? It is almost certainly not a single language model sitting alone and typing out a complete proof from first principles. Language models are stochastic parrots of their training distributions. They are excellent at generating plausible continuations. They are not, by architecture, trusted to perform infinite chains of error-free deduction. The plausible routes are hybrid.
The first plausible route is an LLM plus a formal proof assistant. The model proposes lemmas and proof sketches; the assistant checks each step; the model corrects its proposals until the proof object compiles. This is the AlphaGeometry pattern applied to a much harder arena. In this architecture, the AI's role is not oracle but explorer; the authority remains with the mechanical checker at the end of the pipeline. The second route is an LLM plus symbolic computation, where a computer algebra system performs the heavy algebraic lifting and the model orchestrates the strategy. The third route is an LLM plus human expert collaboration, where the model narrows the search space and the human closes the final gap. All three routes are valuable, but they are not equivalent in their implications for mathematics, for crypto, or for the labor market.
The fourth route is the one that worries me. The model might have solved a problem that is 'open' only in the sense of being absent from public training data, but which is solvable by a competent graduate student with enough time. The word 'open' is doing a lot of work in this story. A problem can be open in the sense that it appears in no textbook and no published research thread; it can be open in the sense that no known publication states a solution; it can be open in the sense that the solution is unknown to the authors of the benchmark. The last, most promotional sense is naturally the one that appears in the headline. If the 50-problem set is an 'Open Problems' benchmark curated as an extension of FrontierMath, the difficulty distribution might be far below the legendary open problems of mathematics. The Riemann Hypothesis is open. P versus NP is open. A curated set of 50 obscure research questions may be open without being deep, and its resolution may be a significant but narrow engineering achievement rather than a scientific revolution.
Consider what is not disclosed. We do not know how the model performed on the other 47 problems. If it solved three and failed or abstained on the rest, the success profile is a thin signal. We do not know whether the three solutions are a complete proof, a counterexample, or a construction that suggests a future theorem. We do not know whether independent mathematicians have reviewed the output, or whether the only review was the benchmark's automated grader. A grader that checks numeric or symbolic outputs is not a proof system. It cannot distinguish between a valid proof and a lucky guess; it cannot check the structure of an argument. The gap between 'got the right answer' and 'produced a sound proof' is exactly the gap between an AI benchmark result and a verified mathematical advance.
I have built my professional life around the difference between declared results and verifiable ones. When I audited token models in the DeFi summer of 2020, I watched high APYs mask systemic liquidity risks. My essay 'The Yield Trap' argued that the market was pricing yield as if it were alpha, when it was in fact pricing tail risk. That experience taught me to read protocols as economic machines rather than as narratives. The same discipline applies here. An AI benchmark score is not a business. A solved problem is not a product. A press release is not a proof. In an environment with so much missing information, the honest analytical position is a conditional one: if the claim is true, then its first-order effects are in science and tooling, not in token prices. And if the claim is false or wildly exaggerated, the market will eventually discover the difference, at which point the narrative premium will evaporate.
Let me now trace the effects that would actually matter, under the assumption that the claim is true. The first destination of such a breakthrough is not a white paper or a crypto conference. It is the research workflow of every mathematics department, and the formal verification toolchain of every serious software company. There are roughly two hundred thousand people in the world who hold a doctorate in mathematics and who are actively engaged in research. If AI systems can assist with conjecture generation, proof search, and invariant discovery, the rate at which these researchers can explore a hypothesis space increases significantly. I estimate the 'augmentation rate' of such tools at forty to sixty percent within five years: more conjectures tested, more cases checked, more errors caught. The 'substitution rate' is very different. Short-term, it is below twenty percent, because the most valuable work is choosing which problems to ask, judging which directions are fertile, and mapping the conceptual terrain. Machines can help traverse the terrain. They are not yet good at deciding that the terrain itself is worth crossing.
This pattern is not new. AlphaFold did not replace structural biologists; it reoriented them. Protein structure prediction moved from a bottleneck of experimental effort to a bottleneck of functional interpretation. The same reorientation is coming to mathematics. The scarce resource is not computation; it is problem selection. The limiting factor is not the ability to generate proofs; it is the ability to distinguish a proof that advances a field from a proof that merely closes a curiosity. In a world where AI can produce thousands of candidate lemmas, the mathematician becomes a judge, a curator, and a strategist. That is a better job description for many of us, but it is a different skill set from the one the academy currently teaches.
The second destination is the formal verification stack. This is where crypto and AI collide with real force. For years, the industry has used three approaches to secure smart contracts: code audits by human experts, fuzzing and symbolic execution tools, and formal verification for the highest-stakes protocols. The current state is fragmented. Audits are slow, expensive, and prone to the same cognitive biases as any human activity. Fuzzing finds shallow bugs well but cannot prove their absence. Formal verification is rigorous but requires specialists who can translate business logic into proof obligations, and the translation itself is a source of error. In 2022, the industry lost more than three billion dollars to smart contract vulnerabilities, a substantial fraction of which were avoidable with better verification practices. The gap between what we know how to prove and what we actually prove is enormous, and it is measured in dollars, not in theorems.
If AI systems with strong reasoning capabilities mature, the verification pipeline will change from 'finding bugs in code' to 'proving invariants of systems'. Imagine an auditor that does not simply flag suspicious lines but constructs a formal model of a DeFi protocol and mechanically verifies that it cannot be drained, cannot be reentered, cannot have its state corrupted by any sequence of transactions. That is the promise of AI-assisted formal verification. The technology to construct such models exists in fragments today. Lean and Coq are powerful enough to represent real protocols. The bottleneck is the cost of producing the proofs. An AI that reduces that cost by an order of magnitude changes the economics of security, and therefore changes the design envelope of what can be built on chain. Deeper and more complex protocols become viable because their invariants can be certified mechanically.
I have watched this industry ship protocols with bugs that a formal layer would have caught, not because the bugs were deep, but because the verification was not done at all. In a market where time-to-market beats time-to-proof, the decision to skip formal verification is rational for the individual actor and catastrophic for the collective. This is a classic tragedy of the commons, and it persists because security claims are not easily auditable by end users. An AI-that-proves does not solve the coordination problem by itself. It lowers the cost of proof, but adoption still requires regulation, insurance, and market incentives to reward certified protocols over uncertified ones. The math may be ready before the market structure.
The third destination is the educational system. If AI can solve research-level problems, it can certainly solve homework. Every mathematics course, from calculus to graduate topology, faces an assessment crisis. Traditional exams assume that a student can be isolated from tools that solve the problems for them. That assumption has already eroded; it will be completely unsustainable within a generation. The response will not be to ban tools. It will be to redefine what is being assessed: not the ability to compute, but the ability to formulate, to decompose, to evaluate the plausibility of machine-generated outputs, and to integrate those outputs into a larger argument. This is a profound shift in the credentialing infrastructure of the technical world, and it has direct implications for crypto. The industry's talent pipeline, its audit workforce, and its research community all emerge from the same educational system. Change the assessment, and you change the skill composition of the entire sector.
There is a parallel here with crypto's regulatory journey that I find difficult to ignore. From 2022 to 2026, the industry moved from a posture of rebellion to a posture of compliance. The approval of spot Bitcoin ETFs in 2024 was not a technology event; it was a narrative event in which the story shifted from 'digital gold versus the system' to 'digital gold as a system'. I wrote about this shift in my report 'The Boring Boom', arguing that volatility would decrease as institutional narratives standardized around regulatory clarity. The same sociological mechanism applies to AI and education. Institutions under existential threat do not simply revise their rules; they revise their definitions of competence. The question is whether the new definitions preserve the value of human judgment, or whether they reduce humans to raters of machine output. I fear the latter is more likely in the short term, and I consider the preservation of human epistemic agency to be one of the most important design problems of the next decade.
The market's response to this story is already predictable. AI-crypto tokens will fluctuate. The narrative of 'AI agents transacting autonomously' will be reinforced. Speculators will buy the bridge between the AI narrative and the math narrative, expecting the two to merge into a productivity explosion. My view is more conservative. The distance between 'an AI solved a research problem' and 'an AI generated yield on a DeFi position' is enormous. The first requires proof of existence and verified inference. The second requires a reliable, economically incentivized, non-exploitable agent architecture. The two may eventually converge, but the path crosses several chasms: formal verification of agent behavior, oracle reliability, adversarial robustness, and the fundamental indeterminacy of strategy in a dynamic market. A theorem-proving AI is a marvel. A profit-seeking agent is a different species. Treating the former as evidence for the latter is a category error, and markets regularly pay a price for category errors.
Let us now be contrarian, because the most dangerous effect of this story is not the overvaluation of AI tokens. The most dangerous effect is the lowering of the verification bar itself. We are entering an era in which machine-generated assertions will be delivered with high confidence, wrapped in statistical plausibility, and immune to human audit. The natural human response to an output we cannot verify is to trust the confidence signal instead. That is a cognitively efficient response, and it is precisely wrong. The entire history of mathematics, and of cryptography, teaches us that confidence is not a substitute for proof. A proof that cannot be checked is not a proof; it is an oracle. And an oracle reintroduces the trusted third party that crypto was built to eliminate.
This is the deepest irony. Bitcoin's invention resolves the Byzantine Generals problem by replacing trust in a counterparty with trust in a mathematical protocol. The protocol is open, auditable, and verified by everyone. If AI-generated proofs become so complex that no human can audit them, and if we adopt those proofs as the foundation of security mechanisms, we will have outsourced the verification layer to a black box. The black box may be internally consistent. It may be correct. But it is still a concentration of epistemic power, and concentrated epistemic power is the exact opposite of decentralized trust. In the chaos, look for the invariant: the invariant is that trust is a social process, and surrender of verification privileges is surrender of sovereignty, regardless of how the output is labeled.
I want to pause on the personal dimension, because the recent history of this industry is also a history of emotional exhaustion. The collapse of Terra and Luna in 2022 broke something in me, as it broke something in many people who had carefully reasoned about risk and still watched the entire ecosystem burn through a mechanism they had not fully counted as a failure of human coordination. I spent three weeks in a cabin in Austin, away from the toxic discourse, reconstructing the failures of Celsius and BlockFi not as engineering failures but as narrative failures: the story of decentralization had been used to mask centralized risk. Solitude is the price of clear vision. That lesson applies here. The story of 'AI solving the unsolved' will be used, in the coming months, to mask the absence of verification, the absence of reproducibility, and the absence of the most rudimentary scientific transparency. The market will want to believe. The wise will want to see the proof object.
What would constitute adequate evidence? First, the three problems must be publicly named and contextualized. The mathematical community must agree that they are genuinely open and that their resolution is non-trivial. Second, the proof must be independently verified, either by human experts or by a formal proof assistant, with the proof object published and checkable. Third, the failure profile on the other 47 problems should be disclosed, because that information calibrates the generalization ability of the system. Fourth, the method should be described in sufficient detail to allow replication: the architecture, the training data, the compute budget, and any benchmark-specific tuning. Without these four elements, the claim is a press release, and a press release is not a scientific article. This is not a high bar. It is the standard to which any serious result in any mature discipline is held.
The crypto industry has a particular obligation to hold that bar. We spent years being dismissed by traditional finance as a casino of unsubstantiated claims. We demanded that regulators understand our technology before they judged it. We built tools for transparency that the legacy system lacked. If we now accept an unverified AI claim because it is convenient for a narrative, we are undermining the very culture of verification that gives our industry its legitimacy. The crowd sees a moon; I see a model. The model must be tested against reality, not against headlines.
Let me now lay out what I think the next narratives will be, because understanding narrative cycles is the core of my work as a market analyst. The current cycle is 'AI solves the unsolvable'. The next cycle, within eighteen to twenty-four months, will be 'provable execution'. The winning systems will not be the ones that generate the most impressive standalone results; they will be the ones that integrate verification into the execution pipeline so that every output is accompanied by a certificate. In crypto, this means smart contract systems that auto-generate formal proofs of their own safety, transaction systems that prove the validity of their state transitions, and agent frameworks that prove their actions are within policy bounds before they execute. The market leader will not be the flashiest AI lab. It will be the quiet infrastructure company that builds the certification layer.
I have been quietly adjusting my own fund's exposure in that direction for months. The public conversation is still obsessed with compute scale and model size. The institutional conversation, at least the one I hear from the traditional finance analysts I collaborate with, is increasingly focused on governance and auditability. The convergence of these two conversations will produce a demand for proof infrastructure that does not yet exist in a mature form. Lean is an academic tool with a steep learning curve. Formal verification languages for smart contracts, such as those built on Rust-based frameworks, are powerful but not yet user-friendly. The gap between the power of the proof systems and the usability of those systems is the commercial opportunity. It is a bigger opportunity than any single AI token, because it is a platform on which all future secure applications will be built.
My estimation is that the 'provable execution' narrative will mature in three stages. In the first stage, which is largely present now, formal verification is a specialist service, expensive and slow, applied to the most critical protocols. In the second stage, AI assistants dramatically reduce the cost of producing proofs, making verification a routine part of the development pipeline for commercial applications. In the third stage, verification becomes a property of the platform itself: every transaction, every state transition, every agent action carries a certificate of correctness, and clients can computationally verify that certificate without trusting the provider. The third stage is where the AI-crypto convergence fulfills its deepest promise. It is also the stage most exposed to the risk I described earlier: if the certificates are generated by a black box, the verification is meaningless. The user must be able to check the certificate independently, using cheap, open, auditable code. Otherwise, we have simply substituted one central authority for another.
This brings me to the ethical dimension, which is inseparable from my interest in this topic. I have spent the last year interviewing developers and ethicists about the alignment of AI agents with human values. The most common framing is about safety: ensuring that agents do not cause harm. I want to add a different framing, which is about transparency: ensuring that agents can explain and prove what they did. A safe agent that is unexplainable is a risk we cannot evaluate. A beneficial agent that cannot prove its own correctness is a liability we cannot price. Blockchain's contribution to the AI era is not merely a payment rail for agents; it is a record-keeping and verification rail. The invariants of computation can be written on a blockchain, and the provenance of AI decisions can be anchored there. The combination of AI and blockchain, done right, creates a transparency layer for artificial intelligence that did not previously exist. That is the vision I call the trustless economy: not trustless in the sense of nobody needing to be honest, but trustless in the sense that every actor can verify the claims of every other actor without relying on centralized authorities.
A theorem-proving AI feeding its proof objects into a verifiable ledger is the first concrete instance of that vision. The proof object is public. The verification is cheap. The history of the discovery is immutable. The system is self-auditing. This is a profoundly exciting image, but it is an image, not yet a fact. And the distance between the image and the fact is precisely the distance between the current unverified claim and the four-element evidence standard I described earlier. The distance is closing, but it is not closed, and the market should price the distance rather than the endpoint.
There is another dimension to consider: the geopolitical and regulatory reaction. If AI demonstrably solves research-level mathematical problems, the regulatory conversation shifts significantly. The laws around data, security, and certification of AI systems will intensify. Governments that already distrust open-source software will distrust open-source AI even more. The response to 'AI verifiable mathematics' will be a demand for audited AI systems, which is a demand for exactly the kind of infrastructure that crypto can provide. But the regulatory instinct will be toward centralization: certified models, licensed infrastructure, restricted deployment. The crypto ethos will push in the opposite direction. The outcome of that tension is uncertain. My experience with SEC enforcement over the past five years suggests that regulators prefer clarity over chaos, and that they will seize any tool that gives them transparent audit trails. The 'provable execution' stack is such a tool; it can be used by regulators to verify compliance without exposing business logic. This is the institutional bridge that the token industry has been seeking for a decade.
Let me also address the question of what this means for the practice of mathematics itself, because the meta-level effects are more profound than the direct effects. Mathematics has always been a discipline of solitude and insight, a practice of deep concentration on hard abstractions. The romantic image of the mathematician in an attic, scribbling by candlelight, has a functional core: mathematical insight often requires exactly the disconnection from social noise that the INFJ personality craves. If AI becomes a ubiquitous collaborator, that solitude changes. The mathematician interacts with a system that is simultaneously a partner and a mirror of the collective intelligence embedded in its training data. The process of conjecture and refutation accelerates, but the experience of discovery may become less individual. I believe this is an aesthetic and cultural loss as much as a practical gain. It is also an opportunity, because the human role shifts toward what machines are least capable of: sensing which problem is worth solving, which question is aligned with human flourishing, which avenue resonates with the deeper structures of reality. The ethicist becomes a co-equal partner with the mathematician. That is a change I can endorse, even as I mourn the solitary attic.
The economic model of mathematics research will also shift. Currently, research funding is distributed through peer review, which is slow, conservative, and prone to groupthink. If AI can generate thousands of proofs per week, the bottleneck moves to the evaluation of those proofs: which are important, which are correct, which open new directions. Peer review by humans will be overwhelmed. New structures will emerge: proof marketplaces, where AI systems submit lemmas and humans evaluate the significance; open formal repositories where the community votes on the acceptance of formal proofs; and certification agencies that maintain standards for verified mathematics. These structures will look more like crypto markets than like academic journals. They will have tokens, incentives, and economic mechanisms for rewarding contributions. The question is whether those mechanisms will be designed openly or enclosed in proprietary platforms. If history is any guide, the open designs will win in the long run, but only after significant contestation.
There is a risk that the current excitement over AI mathematics leads to a misallocation of talent. Every ambitious graduate student wants to work on the frontier of hot topics. If the frontier appears to be 'AI-assisted proof', the field will flood with engineers and statisticians, and the slower human practice of conceptual mathematics may be neglected. That would be a tragedy, because the conceptual work is exactly the work that AI cannot yet seed. I believe the next generation of mathematical breakthroughs will come from a symbiosis: a human who thinks slowly and deeply, coupled with a machine that searches fast and broadly. The human sets the direction; the machine fills the terrain. Both are essential. Neither can be safely ignored.
In terms of my own investment thesis, I am watching for signs of that symbiosis in the crypto-adjacent world. I am looking for teams that treat formal verification not as a marketing label but as an engineering discipline. I am looking for protocols that publish proof objects on-chain. I am looking for AI infrastructure companies that open-source their verification layers rather than protecting them as trade secrets. The open-source signals are the strongest signals, because they align with the verification culture of the best mathematics. A closed verification system is a contradiction in terms. If you cannot audit the verifier, you have not solved the trust problem; you have displaced it.
The current story provides a useful stress test. The absence of disclosure is itself informative. It tells me that the actor behind the claim, whoever it is, believes that the market rewards narrative before evidence. That belief is rational in the short term and catastrophic in the long term. Every actor who exploits the credibility gap accelerates the eventual buildup of skepticism, and skepticism is the currency of our industry. In a perverse way, I am grateful for this event, because it forces a clear-eyed discussion about the epistemology of AI results. It forces teams, including mine, to formalize their own standards of evidence. It forces the market to ask, repeatedly, the question that should be asked of every claim: where is the proof?
I want to return now to the benchmark itself, because the construction of the test set reveals a great deal about the intent behind the claim. A 50-problem 'Open Problems' subset of FrontierMath is, if it exists, a deliberate curation. Someone chose fifty problems that are open, presumably in some sense, and presumably representative of something. The choice of which fifty is an editorial decision with enormous influence on the measured outcome. If the fifty are chosen to be solvable by extending existing techniques by a moderate factor, then a 6% solve rate is an engineering achievement but not a mathematical revolution. If the fifty include problems that have resisted specialist attacks for decades, a 6% solve rate is a seismic event. Without knowing the contents of the set, the result is unintepretable. The difference between those two scenarios is the difference between a milepost and a singularity, and the current reporting collapses that difference entirely.
The secondary effect on the AI industry itself is also worth analyzing. Benchmark scores function as a market for reputational tokens. A high score on a prestigious benchmark justifies future rounds of funding, attracts talent, and shifts the narrative of model capability. If the scoring process is transparent and the benchmarks are well-constructed, this is a healthy dynamic. If the benchmarks are opaque or easily gamed, the dynamic is corrosive. FrontierMath is a serious attempt to construct a robust benchmark, but the 'Open Problems' extension, if the reporting is accurate, appears to have been presented without the methodological transparency that makes a benchmark trustworthy. The AI community will react to this by demanding more disclosure, and the demand will be healthy.
There is a philosophical point that I cannot avoid, because it is central to my own sense of meaning in this work. The question of whether AI can 'solve' mathematics is entangled with the question of what mathematics is. If mathematics is a set of formal consequences of axioms, then a machine that mechanically explores the consequence space is a perfectly good mathematician, albeit a slow one. If mathematics is the human activity of finding patterns, making analogies, and seeking aesthetic elegance, then the machine is a tool, not a colleague. I hold the latter view, not because I am romantic, but because I believe the problem-finding capacity of humans is the scarce resource. The machine can compress the distance from conjecture to theorem. It cannot compress the distance from human curiosity to a meaningful conjecture. The important direction of travel is from human values to machine verification, not from machine outputs to human acceptance.
The regulatory community will eventually need a vocabulary for these distinctions. My conversations with compliance officers at token funds have taught me that the word 'decentralized' irritates them, because it is used imprecisely to evade responsibility. They will feel the same about the word 'solved'. The measured response is to require operational definitions: solved means that a specific proof object exists, has been checked, and is published. Anything less is progress, and progress is good, but progress is not proof. When regulators ask for definitions this precise, they do it not out of pedantry but self-protection. The industry should embrace that precision, because it separates the serious from the superficial.
Let me now turn to the specific question of what the next twelve months will look like, in my judgment. First, we will see a wave of incremental AI mathematics papers, each claiming to be the next step beyond this result, each with varying levels of verification. The signal will be noisy. Second, we will see a consolidation of formal proof infrastructure. The major AI labs will invest in Lean integration, because defining the formal proof object is the only way to make claims auditable. Third, we will see a wave of compliance demand from institutional clients. They will ask AI vendors: how do you verify your outputs? how can I audit your reasoning? The vendors that can answer with a proof object will win the enterprise market. The ones that answer with a marketing website will be filtered out. Fourth, in crypto specifically, we will see a new token narrative: proof infrastructure. The winners will not be the AI models themselves but the platforms that allow anyone to verify the model's work on-chain.
I am positioning quietly. The market is still shouting about the top-layer narrative, while the second-layer infrastructure is being built by small teams with little attention. That is where the value is concentrated. The crowd sees a moon; I see a model, and the model tells me that verification infrastructure has positive network effects, high switching costs, and durable defensibility. The buzzword-driven layer of AI tokens will be volatile and difficult to trade. The verification layer is boring, slow, and deeply necessary. It is exactly the kind of investment that my own experience has trained me to identify: structurally sound, narratively delayed, and essential to the long-term functioning of the system.
I should also address a darker possibility. The claim 'AI solved three open problems' might be deliberately exaggerated to influence policy or markets. Regulation-by-enforcement has taught this industry that a headline can move more capital than a whitepaper. If the claim is false, the eventual collapse of the narrative will produce a crisis of confidence, not just in this specific result, but in the entire genre of AI-verified claims. The old saying is that a lie can circle the earth while the truth is putting on its shoes. In our world, a false proof can dump a hundred million tokens while the counterexample is being formalized. The best defense is a culture of radical transparency: publish the proof, publish the verification, publish the failure profile. The current story, with its total absence of disclosure, is a reminder of how far the culture has to travel.
My personal conclusion is that the story of AI mathematics is real in its potential, but the current claim is unverified, and the current narrative is inflating that potential far beyond the evidence. I remain deeply optimistic about the direction of travel. I believe the integration of formal verification, blockchain transparency, and machine-generated proof will reshape how we establish trust in both mathematics and software. I believe the ethical vision of algorithmic empathy, where machines and humans collaborate on solving the deepest problems without surrendering the capacity to question the machines, is within reach. But the distance from a headline to that vision is measured in proof objects, and proof objects are not produced by press releases.
In the end, the question before us is not whether AI can solve open problems. The question is whether we still know how to verify. Verification is a muscle. It atrophies when we outsource it. It strengthens when we exercise it. The blockchain industry was born from the exercise of that muscle. The next decade will test whether we remain willing to do the exercise, or whether we will surrender to the comfortable authority of a machine that asserts and never proves. The choice is ours, and the choice is reflected in every unexamined headline we accept.
The next narrative will not be 'AI did it'. It will be 'show me the proof object'. Position for that shift. The infrastructure that emerges from it will outlive every token that was issued on the back of the current hype. And when the next crisis of trust arrives, as it will, the teams that built on verification will be the ones still standing.
Solitude is the price of clear vision. The market will reward the people who saw through the headline and understood that math does not care about conviction, that narratives are liquid, and that the invariant we must preserve is our own capacity to verify. The crowd sees a moon; I see a model. The model says: prove it, or be quiet. That is the discipline we need.
Coding the future, one block at a time, is not only about writing code. It is about writing proof, about anchoring truth in a form that cannot be corrupted, and about building systems where trust is not a story but a certificate. That is the work. It is slower than the hype cycle. It is less glamorous than the press release. It is the work that lasts.