# Verified AI: Scaling Brilliance Through Formal Verification

**Podcast:** Latent Space: The AI Engineer Podcast
**Published:** 2026-06-03

## Transcript

But it's for the first time now I think verified AI is to open up collaboration.
Either it's human-AI collaboration.
Well, before blueprinting, that's human-human collaboration.
And lean was a grounding, was a verification, formal language.
And then human-AI collaboration, like we're seeing now, future AI agent-agent-agent-like collaboration.
Like I think verified AI is for openness.
It's not for meeting the requirements of closed industries.
And I think, just like I think verification should not be about, oh, I remember, like, you know, there's this article, like, chatbots mixed up.
Oh, there's a mess solution to hallucination.
Verification to me is not about lousiness.
Verification to me is about scaling brilliance, compounding brilliance.
It's like just kind of going back to the collaboration point.
It's about Ramanujan being a much stronger mathematician.
He was already a really strong one, but verification helps him extend the brilliance, like both kind of scale up and scale out.
Welcome to the Latent Space AR for Science podcast.
I'm Brandon Anderson.
I build RNA therapeutics at Atomic AI.
And I'm joined by RJ Haneke, the CTO of Mira Omics, working on spatial transcriptomics.
It's a pleasure to have Karina Hong, CEO and founder of Axiom Math.
Axiom has made a splash in several different areas.
First, they got a perfect score on the Putnam last December, I think.
They also had the claim of the first AI to prove research conjectures using formal verification.
And very exciting, they just yesterday announced quite a large Series A.
Yeah, welcome to the show.
Thank you for having me.
You just raised $200 million, which as one of your colleagues said, this is like basically the entire like U.S.
math budget for math research each year.
Is that true, actually?
According to his LinkedIn post, yeah.
Okay, wow.
$250 million is our apparently annual math budget.
You seem to spend more on math research.
Yeah, it's kind of sad.
Yeah, I know.
But anyway, like, you know, as a nerd who loves math, that's like really cool.
But I mean, I'm just like, that kind of blew my mind.
Like, what?
Like, okay, so like, yeah, how is it 200 million, I guess 1.6 billion valuation?
Yeah, I don't know.
Yeah.
Well, super, super excited to be here.
Also, I think like, you know, this is a series A, so it's a very, very interesting, timely, timely podcast.
We are like a seven, eight months old company.
So it definitely means a lot to us.
It's a really cool milestone.
We're currently about like 30 people now.
So kind of going into, I think this amount of funding will...
like give us a feel that it needs to accelerate the strong execution momentum that we have so far.
I think like people think of us, like there are many kind of ways to think about Axie and people think of us as a math startup.
So math startup, lean startup.
The other obviously things that we do that are formal verification, we think verification is a really good best first market format.
And so I think this fundraise is going to like let us explore some of the applied domains as my colleague CTO.
Shubo said in the launch video of the Series A we had, it lets us broaden our dreams.
So, yeah.
But still, like $200 million and I guess a $1.6 billion valuation.
How is there a market for that?
I mean, obviously, you're not doing this just for the fun of proving things, although I'm sure there's a lot of that.
So let's bring us back to 2024.
So when, you know, 01 recently models just came out.
What was Anthropic kind of like secretly working on back then?
It was coding.
And everyone knows they're working on coding.
Like OpenAI, Meta, Accent.
Everyone has full knowledge that Anthropic was working on coding.
And they just like overlooked it.
They thought.
Oh, they are at B2B plays.
They just want one vertical.
If you think of coding as one vertical, and now look at where we are today.
Coding kind of like strong transfer learning from coding to reasoning to basically, you know, a monopoly in the future of reasoning.
And I think that's really, really shocking.
The people who are working on coding, I think back then, believe in something that we believe, you know, similarly with math and lean now, which is that if you have more structured and formal data, it's going to be a lot more horizontal than the specific vertical we are tackling.
So, you know, if today we are doing, you know, math the informal way, like the standard chain of thought data, train a math model based on human preference, then I would say, well, perhaps we're just a math startup, right?
But, you know, well.
we are pursuing math, we're also doing things that do have transfer learning to other domains.
So I think that's kind of like the broader picture is that while the DNA of the company remains math and all of us are math nerds, and this is a very strong cultural statement, everyone has a great mission of having AI be a superhuman mathematician like we're seeing on a batch of research conjectures.
In fact, we have another batch coming.
We're also thinking that this is going to be fundamental to verified reasoning.
And we kind of talk a little bit about verified AI.
I want to talk a little bit about verified AI next because I think you have another.
Yeah.
Yeah.
I have several things I want.
So I want to hear about the verified AI.
I do want to dig in a little bit.
So do we know that, you know, Anthropic and OpenAI and everyone, they're not doing formal verification and using that for their rollouts and whatever?
I think I have a lot of, like, rumor mill that I probably shouldn't, like, put it on the record.
Like, I think, you know, like, researchers talk, they play card games.
Yeah.
But there are really interesting reasons if they are or are not doing it.
I think that's, like, kind of the takeaway I have, which is that if you're, like, at a frontier lab and the direction actually does change a lot for a lot of the reasons beyond your control.
So I want to kind of like bring us back to the Alpha Proof moment, right?
Like Alpha Proof was such an amazing, that really the 2024, 28 out of 42 performance was the IMO moment for me.
It was not gold in 2025 because...
Across 2024 and 2025, AI models could solve all the problems that are not combinatorics.
The only difference is that, you know, if you get all the problems that are not combinatorics, you get 28 in 2024 and 35 in 2025 because there's only one combinatorics question in 2025.
After AlphaProof, kind of like we didn't see a lot of the formal math.
you know, results or kind of progress from Google DeepMind.
And that's actually because of reasons that are not necessarily technical.
But if you're at a startup and you have very singular focus that is formal math and verified AI, then, you know, you get to work on a really cool problem for a long time.
And you have like a lot, a lot higher likelihood to get to where you want to be in terms of like progress and breakthrough unlock.
So, yeah, just define that for us.
Yeah.
Like a lot of people think about formal verification as an ancient, you know, subject.
It existed like as long as, you know, way before like deep learning and existed in the time of rule-based computer science.
There's this really strong push of like formal verification around like since ever since 1980s.
Really interesting historic anecdotes such as I think the Paris Trade Union demanded that the automatic switching.
of the subway system needs to be formally verified for safety purposes.
So quite interesting trade union for technology.
And like, I think around the time of Challenger, both before and after European Space Agency was using formal verification for the Ariane spacecraft.
It's also interesting.
Boeing, Airbus for verification.
And then more recent years, right?
I think like there's a lot of push about automated reasoning at AWS because they have a lot of enterprise customers that really require things to be 100%, you know, verified.
And there's no edge cases missed and like just general like testing doesn't satisfy the need.
So a lot of people think about verification as...
something that's like annoying because it's like tax and compliance.
Like it's making sure that we are good to go, right?
And like, that's really not the way.
And so we talked about like verification.
I think our competitor, when they launched, they talk about formal verification, pre-reasoning.
They talked about it in the time of hallucination.
And maybe for them, like formal verification is about the lousiness, the hallucination.
For us, no.
Like for us, verified AI, is about the brilliance.
It's about scaling and compounding superintelligence.
So this is quite a deep point, and sometimes it takes a little bit of explanation.
So if you think about, like, you know, the place of brilliance, for example, Ramanujan, like he's a brilliant mathematician.
He was able to find a lot of, like, interesting formulas just by intuition before he knew how to do proofs.
So he went to Cambridge, you know.
work with Hardy and Littlewood.
And, you know, in the famous movie, The Man Who Knew Infinity, there's this like storyline of how hard it was for Hardy to force him to no longer rely on intuitions and do proofs.
After he learned proof writing, he came out as a much more powerful mathematician whose results, like intuitions, turn into theorems.
And future generations of mathematicians build on those theorems.
So it is a way to kind of scale and compound the intelligence that we already have.
Another example, mathematicians kind of have been writing code in English or their respective countries' natural language for thousands of years.
And why do I call it writing code?
Because there's this sort of community standard of rigorous logical deduction.
Everything has to be step-by-step correct.
Otherwise, you will get outcasted by your math community.
So it's a law.
Well.
Worth rules in the community.
So it's interesting, right?
Because that is kind of human mathematician enforced, right?
And so it's a peer review process.
Peer review of a paper currently takes two years.
Okay, so proof assistants and, you know, formal proof trackers like Lean still found its place.
Right.
And why?
If I'm a mathematician and, you know, my work can be peer reviewed by other humans, like why do we even, why do mathematicians even play with lean?
Right.
And why do we even like talk about kind of like, you know, lean based assisted like serum proving?
It's because like it handles a low level.
For example, we're not even talking about AI.
We're talking about, for example, the grind tactic.
It can currently handle a lot of mass proofs, like at a very low level.
And this is pretty shocking because I have seen, you know, actually another company working in the same space, like, you know, some of their demo.
And I look at the demo, like it can actually completely be handled by grind, which is a tactic in Lean.
Can you explain what Lean is to non-experts?
Oh, okay.
Yeah, yeah.
I think our order is like a little wrong.
Yeah.
So Lean is a computer program, a bit like for math proofs.
It is a formal language, just like its cousin, Isabel, COQ, or ROC.
some other further cousins like Daphne, Agda, like these formal languages, whole sector.
And what does it do?
It basically, if you have a proof written in the programming lane, and then assuming there's not any weird things happening, like unintended use of sorry, which is a tactic that lets you take things for granted, assuming everything is safe, hence people have...
tools like Comparator, SaveVerify, and Axiom recently wrote out Verify Proof, that's like 100 times faster than Comparator, then, you know, once you kind of execute that program, like once it compiles and it tells you that it's correct, then the proof is actually correct.
So it's like a type checker?
Yeah, that is based on this result called Howard Correspondence, which turns proofs into programs.
So I want to talk about the magic of Lean.
Why I think it's a really good programming language is because...
On one hand, if you don't care about the formal part at all, if you don't care about the logic part, you just want to use Ling to write code.
You can.
Like we have had candidates actually currently, you know, the person is working at the Ling FRO.
He wrote Autograd in Ling in our interview process.
So it's a Turing Complete language.
That's right.
So you can write, you can do a lot of things with Ling.
It's a functional programming language.
Okay.
Right.
And then you can also use it to, so you use it to do.
code and you can use it to do math, two-in-one.
Okay.
And kind of going back to what I was kind of getting at, if mathematicians are...
already enforcing that most proofs, you know, say maybe not all mathematicians, but the ivory tower people in academia, all proofs are correct.
Why do we even need Lean, the model tracker?
It's because Lean has tactics that help them handle the low-level calculation or proof or deduction, not calculation, then for them to be able to navigate in a high-level intuition space.
So this is my point, that it is not about like formal verification or verified AI to us, it's not just about handling or like kicking out the lousiness, the hallucinations, the mistakes.
It's about scaling brilliance.
It's about super intelligence.
I actually, Terrence Tao has a great video also about using lean to, as a way you can collaborate.
That's right.
Because you can use the sorry.
Exactly.
That's another point I want to talk about, right?
A lot of people think about, you know, what is our market?
It has to be some like really niche industrial societies area that is mission critical, safety critical.
No, that's not the TEM.
The TEM is all code.
The TEM is a...
It's a right of first refusal on all AI-generated code.
Like right of first refusal, meaning, you know, you get to choose whether you want to verify it.
So this is the important part I want to kind of come across, which is that people talk about formal verification as...
almost like painful because it has all these like stringent requirements.
Up until now, it has been.
Yes, yes.
And to us, it's actually verified generation means performance gain.
It means higher sample efficiency.
It means a startup like us with like, you know, still we raise some money, but lesser compute budget, lesser data budget, then Frontier Lab will be able to match and even exceed, you know, performance on superhuman tasks.
In fact, for the pandemic exam that we just competed, December 2025, which we did in real time, Mass Arena, which is this organization that evaluates a lot of LLMs, found the best LLM, DeepSeek, got 103 points out of a 120-point exam.
The best human, obviously, we now know is a student from either MIT or Chicago.
We don't know which one because they don't announce the top five winner score, got 110 and we got 120.
So it's the first time, actually, I remember when we were starting this, people were like, is it even possible?
that a formal mass, you know, system with so much orders of magnitude, less data can match or beat an informal LLM.
And PUNEM is the first time it beat, right?
And so we're not thinking about it just about the painfulness, the challenges it pose.
We're thinking about the verified generation performance gain, the improvement, the fact that you can, you know, just like you would expect.
RL4 lean to have improvement because of seeing evidence of RL encoding.
So this is the second point I want to make about how to think about verification, verified AI.
So maybe we can talk a little bit about why.
Can you describe what is different about what you do versus what the Frontier Labs, you know, at least when they're building their standard RL enhanced LLMs, what...
What's different about what you do?
Yeah.
So we heavily rely on kind of data called Lean data.
And we kind of talked about Lean is all the data that we have that's Lean proofs, you know, it's correct.
So you know it's correct or not.
And that's quite important.
So, you know, we have a system of models.
These models are post-trained and using RL or SFT.
So LLMs.
like some sort of foundation model that you get off the shelf and you post-train it or can continuous.
Yeah, and there's obviously an inclination for open source, you know, base models.
So it does speak English, probably knows how to code, but also you fine tune it or continue.
Yeah, and the base model might be similar to what everyone else is using as well, right?
If they're not kind of pre-training their model.
Right.
And then we basically do this, you know, RL for formal math kind of.
There's, I think, a standard pipeline or like, you know, tricks of the trade that people use.
We try to innovate really on top of it as much as we can.
I think that we found scaling inference to have almost no wall, recursively decomposing, you know, a proof goal into many sub goals and learning to backtrack as well.
Is there a risk that like you start out with this, you know, what you know in a certain domain of data sets and so on, and then you start rolling out, you know, recursively in a space, but now all of your training data is localized in some domain that you still is only so like maybe logarithmically.
in some large space growing from your initial training data.
So you could get trapped essentially in that, you know, you could be really good at this, but you just created a big jagged frontier where some other domains are just far from them.
Distribution shift.
Yeah.
Are we talking about?
So, yeah.
So, so, you know, it is an open question whether a, a system that can do really well in number theory can do well in.
Give me, you know, another field of math.
Yeah, exactly.
Well, actually, I think the way we think about it is it depends.
It depends on whether topology has a lot of the existing definitions as almost like, you know, the math infrastructure existing.
Because what people have found in the past is when people were building out MathLib, like, you know, for the algebra, you know.
book work.
Like they can just.
So Mathlib being the lean like undergraduate library.
That's right.
So it's like all the proofs that you learn in undergraduate math.
Yeah.
And they're all sort of in lean.
Yeah.
So for example, some of my friends who currently are at Axiom is, you know.
crazy like full circle back moment.
Kenny, we're like friends for like, you know, five, six years.
And he was the first one to tell me about Lean.
He was working with Kevin Buzzard to build out MathLib.
It's a lot easier to codify algebra in MathLib than for analysis.
So that's interesting because for analysis, a lot of the definitions around convergence limits, et cetera, becomes tricky.
And so I don't think there's a lot of like topology in mathlib today in terms of like differential topology, differential geometry kind of stuff.
So, you know, our system likely will not do very well on those domains because it doesn't even have definitions to build off on top of.
places where the definitions are in, we actually are doing quite okay in terms of distribution diversity.
We have good performance, you know, having solved open research questions and number theory, commutative algebra, algebraic geometry, some discrete math that come into our exam probability.
So earlier you said that like with the Putnam exam, the 2024 version, when all of the questions were that were not that alpha proof did not get right.
The IMO, International Resolute.
Yes, for the IMO, all of the ones they got wrong were in combinatorics.
Is there like a weakness there in that specific domain?
I would say so.
For Olympia in math, people are seeing combinatorics being a little bit more tricky.
Seems like the steps are quite creative.
So I'm a human and, you know, when I have...
friends who are really good at combinatorics, which I never consider myself really the top of combinatorics.
I'm kind of better at number theory, but I know some people who are just, they're IMO gold, perfect score, Putnam fellow, perfect score, and like all the way.
And then when they do like tricks in combinatorics, I'm like, I don't know how you thought of that.
But, you know, after you give me that construction, it actually becomes a lot more trackable.
I think a lean-based system will struggle in those very creative places, which is why we at Axiom actually also invest on something called mathematical discovery.
It does not use lean.
And we have some major news in the coming.
basically open sourcing entire code bases of mathematical discovery coming up.
You want to tell us a little bit?
Yeah, yeah, sure.
So we are currently having two code bases being open sourced.
The goal is for if you're a mathematician or you're a theoretical physicist and you have a problem that you would like to solve, for example, you want to find a construction that is a very complicated graph construction, then we would suggest you follow the very detailed.
manual supposed intended for mathematicians to run the code that we write.
It's a tool for mathematicians to make mathematical discoveries.
Mathematical discoveries is this idea that, you know, proof is not enough.
for math.
In fact, before you kind of start proving something, you don't know where you want to start.
So you will try to construct some interesting examples.
These can be usually, say, sequences, right?
If you want to understand the property of a sequence, you will write out a few of the first terms.
This can also be graph.
So if you want to, you know, figure out what the graph that you're looking for.
should I have, say, a certain property, then you will start by doing some simpler version of the graph.
Now, constructions cannot be done by Lean.
So we believe in having AI for mass discovery.
And we have, you know...
One of the OGs in that field, François Charton, a member of technical staff at Axiom, and he previously had done patent boost and end-to-end, you know, set out disproof, a 30-year-old conjecture by finding a counterexample, found the solution to a 130-year-old problem, the global lampana function, that is a kind of mathematical object showing up in the three-body problem.
So we are thinking that, you know, mathematical discovery tools should be open to the mass community.
open-source, the entire codebases for that.
So discovery meaning it makes new conjectures?
Yeah, it's a pre-conjecturing step, actually.
Okay, oh, I see.
Yeah, so you start to form intuitions.
If you're a mathematician and your goal is to solve a really hard, hard conjecture, Ximprover can't just solve it for you.
You might want to...
try to formulate some sort of lemmas, conjectures that you want to say then give to Axiom Prover.
If you're a human mathematician, you will start by wanting to formulate that conjecture.
You don't know where to go.
You want to find constructions.
Now, the code basis that we're going to open source is going to help you hopefully significantly.
So one thing that maybe there's a lot of computer scientists listening.
And one of the things that will immediately kind of come up, especially when you're talking about formal verification and so forth, is Rice's theorem and decidability and incompleteness theorem and maybe some arguments about computational complexity in LLMs.
So I'm curious to hear, Rice's theorem says you cannot prove non-trivial things about programs for all programs, right?
So how are you navigating this space?
Obviously, formal verification, you know, is able to do some things.
Yeah.
So, yeah, I think, like, it's very clear that you just, like, there's the radical result telling you you cannot formally verify all programs, right?
But I think it's good to formally verify majority of the useful programs, right?
So, you know, like, I remember there's this MIT, like, little, like...
documentary or not a documentary, like an advertisement for, you know, people who are admitted students.
And then there's this famous line by Tim, the beaver, the mascot of MIT, saying that, what does theory give you?
Which is kind of like, it doesn't stop us from trying to push it as much as possible.
So the goal that we have for the future is suppose you are, you know, doing...
doing the coding, you want to web code a really complex task.
So, you know, currently it's front-end websites, but in the future, we might want to web code much more complicated things, whole distributed systems even.
Then we want to be able to, say, decompose it.
There's maybe a high-level kind of like sketch plan this we can make, other people can make.
But say, you know, you have clock if you like, you know, kind of break it down into 10 things.
And at one point, it will decide to call Axios.
And Axiom will give you a computer program that you know is formally verified.
Or it will say, this is still too hard for us.
So you write the program?
Yeah.
You give it to Axiom?
Yeah.
It makes changes to it, maybe?
So we're talking about kind of two sort of phases.
it is possible that we are the verification partner.
So you already have a computer program and you want us to verify it.
In fact, like, you know, GPT found a proof to an unsolved ERDOH problem and our competitor Harmonic, you know, Aristotle, you know, verified it.
But we can do, we want to do verify generation, right?
We might want to say, hey, you know, this little component, everything that we generate and provide for you is formally verified.
I see.
So the idea would be you co-generate both.
And I can imagine this fitting into the idea of a promise or a sorry, sorry, and then a sorry.
Which is mean sorry.
Meaning it's a lemma that is unproven, but you're just taking it as given until you can have the time to prove it.
Is that a good way to think about a sorry?
That is a good way to think about a story, but not necessarily in the coding context.
So I can imagine you can say, assuming that this module is verified, then this module is correct.
And so that you can decompose a problem small enough that you can verify?
Is this kind of the intuition here?
So let's say we want to, you know, like web code control flows.
Yeah.
Right, that's quite...
hard, you will likely, you know, break that down into multiple steps.
And then it will continue to break down these steps into more fine-grained steps.
And at one point, you want something that is absolutely correct.
And then this is also something that is likely within reach.
Then we want to generate, you know, both.
We want to generate a piece of computer program.
And underlying is a guarantee.
that there is also the proof that has been generated, which tells you that the thing that you specify, this, you know, program can solve for you.
So the vision we have is anything that can be, which anything is, you know, and it's a little bit marketing because, as you said, the theoretical abound, but mostly, well, almost surely, hopefully, anything that can be defined can be executed.
Anything that can be specified can be proven.
So the way I think about it is if you have a program times a, you know, a program times a statement or problem, it maps to verifiability conditions times a proof.
So while the program verification community has...
given you, say, the verifiability conditions, and we're trying to kind of recruit a really strong team to help us do that, Action Prover is going to give you the proof.
So just help me map from the program to the proof.
Because I could say, you know, this two-line lean program verifies, you know, sort of like whatever.
whatever I claim it solves, how do I know that it actually verifies the thing that I think it verifies?
Yeah, so for example, there's this benchmark called Verena.
It's a code verification benchmark that's supposed to be limb-friendly.
And so, you know, every problem is a...
coding problem.
And the goal is to generate, there's a code part and there's a proof part, two different computer programs.
And then the goal is to generate code with proof.
So, you know, the code that supposedly solved this problem and then to prove that this program indeed.
does solve the problem.
I see.
Now, how do people do on this benchmark?
I kind of want to like talk about this a little bit because it's interesting.
It was rolled out, I think, by Berkeley and Meta researchers in 2025.
And they found, I think, whatever version of GPT they evaluated does like pass one, like 3.6%.
iterative, something like 22%.
Now, you know, how does the formal mass systems models do?
Cobra, which is a system, because in a system you iterate and define, so pass one doesn't quite work, but still they evaluate it.
Pass one of the system, about like, I think 11%, 12%.
And then also DeepSeek Prover and Godot Prover, very strong Prover model, 11%, 12%.
And I think our competitor has released last year, the only proof part, 96%.
And we actually recently, with no modification to the Putnam system, we saw a 99% out of the 189 problems.
We saw 187.
We missed only two.
Code-wisp-proof.
So if you want to train something to do code-wisp-proof and you want to do reinforcement learning, it's actually quite annoying because, look, it's mixed.
If you want proof to be informal math, it's very annoying because then that's like just mixed objective function.
Your code is something like Python.
Your proof is a natural language math proof.
You will not have a strong RL kind of performance, right?
But if you have proof as Lean and you have, you know, code, you can choose Rust, which is a strongly typed language.
It's more conversion.
So you're going to have much better performance.
I can't wrap my head around how do I tie.
So like I can say that this proof solves Fermat's last theorem, right?
Yeah, yeah, yeah.
I don't know that like.
Yeah, yeah, yeah.
But it's two lines in lean.
Yeah.
Obviously it doesn't.
So how do I know that the program that I wrote matches the proof that I generated?
You will basically look at the coding problem and you look at the program and then you like try to see if it satisfies the verifiability conditions.
But like, how do I know?
Right.
Like if I read it.
Right.
You know, like I can, I can just like eyeball it.
Yeah, yeah, yeah.
And then like traditionally how mathematicians have done this is they, you know, they take the paper and they read it and they say, I agree that this proof.
solves the problem.
And then this other person says, no way it doesn't for, you know, like, look at this and then people disagree.
And eventually there's consensus that, like, this proof solves this problem.
So, like, how do, how are you?
But you check it step by step.
Yeah, right, right.
Yeah.
So you basically will look at the verifiability conditions and see if it does actually satisfy that.
So suppose we're looking at a piece of computer program, right?
And then whether it does actually solve the coding problem, you will have a judgment about that, right?
Yeah.
So you will not solely rely on testing, even though that is the way.
So somebody looks at the proof and says, yeah, that actually solves the problem that we think it's supposed to solve.
But then now you're basically...
producing a, you know, formal verification program that satisfies the verifiability conditions about this program and this statement.
So again, the function is taking you from the program and the statement to verifiability conditions and proof.
Okay, so I can see how this works in a benchmark.
Then if I have, let's say, I have a flight control system that is like very complicated.
Then the problem becomes very annoyingly, you know, the, like, specification.
I think the word is gonna, you know, even if we say successful, like, like, anything that, you know, that we will have a specification problem.
Yeah.
So, like, here comes a bank saying that, like, please, do I have a really safe financial audit?
Sorry, like, prove the financial audit for me, right?
Yeah.
Like, what does that mean?
Like, we can't specify.
Humans are bad at specifying everything that we want.
Right.
There's always, like, some sort of saying that we are not specified.
And if it's not specified, it's not proven.
Okay, so what do you do about that?
Yeah, so we're not there yet.
Okay.
Currently, you know, like, again, the vision as of currently is anything that can be specified can be proven.
Okay, yeah.
Now, obviously, there are people who have been really good at, you know, that's maybe where that's the informal kind of reason I come in.
The informal reasoner can, and this is, I want to kind of, you know, call the literature of testing, like testing are great because testing is like, hey, have you thought about that?
Like I want to highlight the work mutation based, you know, LM unit test generation by XMCTO Shubo and he was a director of Facebook AI Research.
Like the way you kind of think about it is like the AI will be like, hey, have you thought about, have you thought about this?
This case.
And so this is a little bit like conjecture.
So the conjecture is going to help with the specification.
I see.
And then the prover does the proof.
And so this is an interactive process, maybe, that the person, so that when we're actually giving good specifications as well.
Yes, I think this is the future of coding.
And I think this is where, you know, this is where I think even if we are supposed, like given the assumption that everything can be formally verified, you know, like studying sort of like, you know, automatic test generation is still interesting because it is basically giving you the specification proposal.
Yeah.
Right.
And then another thing is.
Let's talk about auto-formalization, right, which is the ability to define it.
It is kind of converting something that is more informal into something that is more formal, auto-formalization.
So suppose I have a coding problem that is written for ICPC, and this problem is written in English, like Alice and Bob, blah, blah, blah.
Okay, now I want to convert that into a formal statement, like a formal spec.
How do I do the auto-formalization step?
Right now, this is going to be because I have not solved the problem.
So I don't have any signal.
I don't have any grounding.
The test cases input output pairs is going to ground my.
So I know I have to know I'm going to give this input.
I'm going to give this output.
It has to have these characteristics.
And so and so I write test cases and I write.
So is there an equivalent in Lean of this, right, where the specification where you just know the sort of like outcomes that you are expecting?
So that you like the statement of the result and then but the proof is completely unproven.
So Lean is actually quite annoying because it's like a lot of the times it's proof.
So you don't actually have the numerical answers to ground it.
Okay.
So auto-formalization is quite a hard thing to do.
Yeah.
Because, you know, what's generally happened is you can't, you just, it's hard to ground the auto-formalization of a statement.
You can obviously ground the auto-formalization of a proof, but because you can then just run it.
But you need human to eyeball it.
How big is a lean proof of like a formalized, you know, of a formalized program of significant size?
I mean, do they grow with the size of the program or do they grow super linearly?
Yeah.
Currently, actually, you know, for each line of COVID, there could be like 20 lines of proof.
Okay.
It's not looking that great.
But is that like a linear relationship or is it as the complexity of the program gets greater than it, like it, you know, sort of also grows?
I don't have a good answer to the scaling law of that.
Oh, okay.
Yeah.
Because I know that that's a problem in formal verification.
That's right.
Where you have these huge, like you have to have these very, very long proofs for even simple programs.
So then are you going to run into sort of like limitations in the capabilities of LLMs when you start to get to larger?
What we believe fundamentally is we are building a reasoning engine.
We have seen Axiom Prover deal with really huge trees that are like, you know, tree of a proof.
We have seen it scale from 40 nodes to 4,000 nodes.
So wait, sorry, Axiom Prover is the LLM?
Axiom Prover is an ensemble system of multiple models that we...
do post-training.
I see.
Okay.
And also it also includes obviously the tools that we have open released.
Sorry.
Yeah.
So we have seen it being able to deal with more and more complex tasks.
I see.
We don't think it's particularly bound.
You could ask, you know, is it bounded at one point on the pre-trained base model?
Yeah.
I think that's a good question.
I think, you know, mid-training could be very interesting because it does actually, you know, a lot of the sort of capability gain does come from that part, right?
If you could argue that even if you try to reinforcement learn some.
person who is not very talented, that person might behave, perform a lot less well than an un-post-trained Ramanujan.
You can argue that way, a very sad reality of things.
So at one point, we might consider doing that.
But we think there's so much to push.
So you just feel like there's so much overhead right now, or so much...
space to glow, that you're not running into theoretical constraints at this point.
I just wonder because, you know, there's been recent results in the computational complexity of the problems that LLMs can solve fundamentally.
And I don't think that they're really a concern for, you know, when I'm writing code with cloud code, but I can imagine.
problems becoming big enough in a system like this where you have a gazillion lines of lean.
You can't get them into the context window, so you have to be smart about that.
And then you have to summarize, and then you're summarizing and summarizing, and pretty soon you're kind of losing track of what's going on.
It just seems like with a very large system like that, you might run into it.
Yeah, I think this is interesting.
It's always a problem of abundance.
So it's important you just keep...
Really, the mathematical discovery renaissance has come.
Action Prover does try to prove everything.
You end up with like tens of thousands of lines of lean proof.
So first of all, auto-informalization is a lot easier than auto-formalization, minus the problem of no grounding, right?
So, you know, every model has seen a lot of text and a lot of lean.
So you can always...
you know, convert that link back into informal.
And then there's the problem of, well, how do you know if you're correct or not?
You can rely on cyclic, like, consistency.
So you then formalize again, like, proof, like, program equivalence, something like that.
So that's...
Oh, so you, like, informalize and then formalize.
Yeah, you can use it to ground.
Yeah, yeah.
Like, and although informalization is, you know, obviously less hard a problem.
So you can always do that.
So for a lot of the, you know...
the link code that we output, we can have an informal summarizer of like big chunks of link.
It's actually doing okay.
So, you know, that's the thing.
And there's another question of like, which I think is very interesting is I think there's a panel at IC.
ML Vancouver last year at the AI for Math workshop.
There's like Leo Damora, Jeremy Avigad, and Shubo and CTOs there.
And they were talking about like, will humans or mathematicians at some point stop trying to understand what's going on there, right?
Because like...
Suppose you're a really ambitious mathematician.
You're like, I want to proofread my hypothesis.
And bang, here's a lean proof.
And like, it's actually correct.
And it's just like, you know, problem, one million lines.
Yeah.
Isn't that like a big negative for the community?
Because, I mean, usually when someone comes with a big proof of something, oftentimes...
Yeah, I was about to get there, right?
It's like, well, will that negative outcome happen was the question the panel was discussing.
It's completely hypothetical.
No one's, like, you know, model system can prove Rima hypothesis, right?
So, disclaimer, please don't cut that part.
Just stand that long.
But, like, you know, will people still try to understand what's going on?
And I think the answer is usually, it's always yes.
I think curiosity and the desire to understand what is going on, you know.
mathematically or in other domains as well, it's a basic human need.
And I think that is like, I think, a dose of optimism in an era of, I think, verifier superintelligence, suppose we get there, is that even if all the outputs are going to be produced and at a much, you know, faster pace and much more exponential volume.
compared to what humans could possibly consume, they're still going to try to consume it.
And they're still going to try to consume the ones that they deem important.
So then basically attention is the bottleneck.
And if attention is the bottleneck, then really intuition and taste, you know, of which statement is probably worth the consumption of human.
And also maybe in a finite compute resource, worth the consumption, worth the...
the sort of spending of compute resources.
That's where human mathematicians' taste will always guide us.
And I think that's incredibly beautiful.
Is it worth like internally taking like results so you can prove one way and then trying to send your system at many different routes to get like orthogonal, conceptually orthogonal proofs?
And so you kind of get a diverse set of different ways of you reasoning about the same thing.
Because, you know, I think it could be very valuable if you give it a problem to say, oh, well, like, here's kind of the brute force natural way that, like, maybe some humans would do it.
And then there's, like, a really much shorter, elegant way of doing it.
So have you essentially thought about training your models to be elegant in some way?
Yeah.
At one point, we're going to get to there because, you know, I think the conjecture will probably depend on what, you know, what.
probably did depend on what we mean by taste, elegance.
Feels like an alignment problem to me, you know?
Like, you know, who gets to say what is elegant?
Humans get to say what is elegant.
That's what I mean.
Human preference.
Right?
There's something about hard work, right?
What you work on hard is what you're going to be good at.
Yeah, yeah.
And we're going to have a problem about that, I think, like pretty much in a lot of the domains as well, right?
Not just math.
Like, how do you be that senior programmer with, you know, really good high-level understanding?
Well, I guess full-stack understanding, high-level and low-level, if you haven't spent the year of training.
I mean, I would argue that you don't, this is very philosophical, but like, you know, I don't need to be good at assembly language programming, right?
Like not many people are good at that.
A few people are because it's important for their job.
So it's not experience, but curiosity.
Yeah.
So, but it feels to me a little different because not being good at like proving things, for example, right?
That seems like a fundamental gap.
in like that maybe my mind doesn't develop in the same way if I am not doing that.
Whereas if I'm just not good at assembly language program well, but I'm good at like higher level programming, so maybe that doesn't matter.
I think that's probably because how the, maybe how the education system, the pipeline works, which is that if you do not show early signs of brilliance, you don't sometimes go through the process of pre-training in math.
Yeah, yeah.
Right, like, so that maybe you can argue that you don't need to, say, you know, learn everything to develop a sense of taste.
But there's, like, a threshold you kind of need to meet.
Yeah.
So, for example, you probably need to be able to code, even if you don't need to understand assembly language.
And that thing might transfer my intuition or, you know, my intuition might transfer from the Olympiad.
mass problems into some other research areas and I tried to pursue and common tharks transfers more direct.
It's very similar and number theory could be further but still okay.
And then when it gets to like something that's a lot more different than Olympian mass, it transfers less strong.
But kind of like, you know, you need to diligent, as you said, right, you need to diligently go through some amount of training.
Yeah.
If people over rely on strong AI, that doesn't happen.
I want to switch gears.
Yeah.
You mentioned software verification.
What are the domains?
How are you going to make enough money to justify the valuation?
And congratulations.
Thank you.
So give us the high level summary of what is the vision that you put in front of investors about why does this actually make a lot of money?
Yeah.
So first of all, this round is kind of preemptive.
So I think a lot of the investors have pretty high interest about Axiom.
In terms of kind of what we believe in, we believe the future of coding is going to be somewhat constrained by verification capability.
And we believe in solving formal math is a very natural starting point.
And then by extension, you can increase the verification capability across hardware and software.
And for hardware, for example, that's quite revolutionary.
I mean, that is, there is no, as we know, there's no partial credit for a mostly verified GPU.
No.
It's all or nothing.
It is all or nothing.
And you do need a perfect prover.
I want to stress this point, which is that suppose I am someone who loves solving math.
I think there are a lot of Twitter users who enjoy Pokemon hunting or those problems.
And then I just try to use a non-deterministic LOM, like GPT say, to try to get the full proof for that.
Now I can do that many, many times.
I might succeed and I might not.
And I might not have a problem with whether I actually succeed or not.
This absolutely does not work for hardware verification.
So for those kind of domains, which I call like hardcore verification needed, it is a pinpoint.
It is a current pinpoint.
There are hundreds of humans and thousands of licenses being dedicated to solve one local grid problem verification.
Just as an aside, my understanding is that the industry standard for design to verification in ASIC project is like one to three, one to four.
One to three, one to four, correct.
Both in, say, team size and in duration.
Yeah.
Right, so if you multiply that, yeah, square.
And then I think, so that's, I would say, like, you know, it's a must cover.
And now for software verification, it is interesting, right?
Because, you know, as probably we all realize, like my nephew.
web codes, a lovable website, there is absolutely no need to formally verify that piece of code.
Like, why would you?
Now, I heard a story from KMAT, actually, that New York Times reporter who told me the story, which is like, However, if you think about like, you know, in the time of agents, like my open claw can probably do all sorts of things and probably can do some bad things.
Like my open claw can decide to like text something bad to my professor, right?
Like, and you can say that.
Perhaps is that a problem of formal verification?
Probably still not, right?
You can change something about the action space and make it more limited.
So you don't need to rely on formal verification.
So you can have a lot of cases, but you can think about, you know, maybe an enterprise that is dealing with a lot of regulatory kind of stuff using agents.
They might want to do something like it's their choice.
But I will argue that the improvement.
of verification capability, both in latency, you know, and inaccuracy, all these stuff, the performance holistically is going to determine whether people rely on formal verification or not.
Sure.
So in a way, we want to make it so good that basically we can make that a choice.
So why did the investors think that you could do this?
Right?
Because, I mean, people have been working on verification for so long, and I think everyone agrees it's an important problem.
And I think certainly if I can just have a verification proof for every program that I write, like, hey, Claude, like, give me the proof also.
And then it just produces it.
And, oh, yep, looks good to me.
I would absolutely do that.
But so why is it, what was it that the investors saw, in your opinion, that persuaded them that, okay, this is the moment I'm going to put in my 200 million or whatever?
I think when it comes to faith, you either have it or you don't.
So either dream the dream with us or you don't.
And that's okay.
Because when we realize the dream, the company is going to be worth 10 billion.
Yeah.
So I think that's kind of the feeling that I have, which is that we believe verification is the critical, critical part to superintelligence.
Our version of superintelligence is absolutely verified.
We don't think there's any other possible future.
We do not believe that, I'm going to say on the record, we do not believe that.
an informal mass system is going to be the mass AGI solution.
Why not?
We just don't believe that.
I mean, the counter argument is, oh, you know, like we just do a lot of good RL and, you know, we've seen GPT, you know, solving, you know, I think some Arrows problem and like whatever.
So why do you think that that runs out of gas?
Yeah, so you can say that if you're on a Frontier Math and you have like, sorry, Frontier Lab and you have like infinite resources.
There's this, by definition, no running out of gas, right?
If you think like infinite means like there's no running out of gas.
I don't think it's going to scale to super intelligence.
So you think that you run out, like you run out of money, basically, you run out of power.
So we as a startup, first of all, cannot do that.
We, first of all, as a startup, cannot do that.
Yeah.
We generally think that formal math and by sort of converting math proofs to programs to code give us much better performance.
So it's just, it's your sample efficiency argument and so forth that you just, and maybe you just can't, that you can't bend the curve enough if you don't use formal.
The thing is, the thing is the informal stuff is also available to us in a way.
If you really, like you can have both informal and formal system and that.
is going to be very strong.
The thing that I kind of like, I think my suspicion about like, you know, whether we can scale to mass AGI just by the informal approaches, you're going to keep having, you know, the LMS judges solution or you have human experts who grade and it's just human experts like doesn't scale that well.
And if you really argue infinite.
infinity, then sure, then you also have infinite money and you can pay infinite.
There's so many.
Is there really infinite number of people who can understand proof at, say, like about like, you know, a result, a non-true result in Langlands program?
I think, you know, good luck finding those.
And in fact, I think how frontier math came together is because they couldn't assemble a benchmark by, they are expert pool, so they have to, you know, collaborate with EPOC to do it, right?
And I think that's kind of what I worry about, about having the human part.
So they have LMS judges and then now stochastic judging.
The problem is like whether something is impossible to achieve versus something is incredibly expensive and like really incredibly expensive and incredibly expensive to achieve get kind of like mixed in the end.
And then, of course, investors always want to know why you.
Right.
So I've read a little bit about your background and I think we would do a disservice.
for the audience, if you didn't hear a little bit just about your personal story.
I see.
Do you want to talk just a little bit about like, you've done some really interesting stuff.
So I'd love to hear like you and then your team.
Yeah.
What makes Axiom special?
Yeah.
I think XM is like very special because they are really expert mathematicians.
Basically, they are users of the system we are developing.
And that iteration loop is very fast.
It is extremely fast.
You have like some of the strongest, you know, mathematicians and both in research and Olympic contest.
And you also have people who are, you know, math loop contributors, maintainers, developers, link gurus, really.
And combine them with people who come from like a...
applied ML, really strong organizations like MetaFair and Golden Age of Fair, as well as people who have cogent expertise, who work with like compilers like Kernogen, have kind of these backgrounds of people together.
I think that's sort of interdisciplinary way of thinking about things quite helpful.
We think AI for math has traditionally been quite interdisciplinary.
People are borrowing techniques from even AI for science.
pure tech, broad borrowing tech techniques from the co-gen literature and pure borrowing techniques from obviously the broader like, you know, frontier like applied ML to try to apply on the niche problem AI for math.
So we also think having this sort of very special, special team is a differentiation.
We also think that.
As you say, there's no permanent mode.
The proprietary data that we generate and a little bit of a flywheel we are seeing is a time mode.
Well, me personally, I love math.
I think, you know, I kind of have been doing math since I was very young.
And like math sometimes gets really hard when the problems you are solving are just a little bit out of reach and it gets a bit depressing.
And times to times I wonder if I can just have an AI help me.
And yeah, I think I figured why not build such a thing.
You did a master's at Oxford in neuroscience.
Yeah.
Has that informed your thinking here?
That's a great question.
I think like my, you know, experience with neuroscience is you learn very well about what's hard and what's impossible.
I mean, it's very interesting.
I think that year of neuroscience like gives me some...
feelings about what's hard and almost no feeling about what might work.
But I think I was kind of under the pretense of neuroscience, like hanging out at the UCL Gatsby Institute and was fortunate to do AI research with some really cool faculties.
And so I think that was a very productive year of AI study, not neural study.
So it was mostly for you studying AI.
That's right.
That's right.
I think in the UK, if you back in the...
you know, 20th century, if you call something AI, you will not get the donation.
But if you call something brain science, you might have a chance.
So the UCL Gatsby, which is a premier AI hub where a lot of people actually go, you know, from their ship to DeepMind, including Demis himself.
It's a very wonderful research environment.
I remember those kind of like tea time talks were very amazing.
And people were basically just doing AI.
It's called the Gatsby Computational Neuroscience Institute.
Yeah, I think how that kind of, you know, happened was because, so I was in the master neuroscience program and then quickly realized that you need to like kill rats and kind of don't want to do that.
And computational neuroscience sounds more appealing.
And when you look at the project and you see like transformer, you're like, you absolutely want to do that.
Yeah, we're all excited about that.
So after the Gatsby, you started a math PhD program at Stanford?
I started actually one year full time at the law school.
Oh.
Because the JD-PhD program is structured in a way where you have to spend one full residency year.
So that was also a very fun year of learning things that are just quite fascinating, like criminal law, looking at homicide cases.
Do you ever feel like the legal system is under or over specified in some way that maybe you could you could maximize and improve?
That's a great question.
I think for a lot of things, it's definitely under specified.
For some other things, I was actually quite excited about sort of transfer learning from mathematical reasoning to those specific fields.
I think appellate litigation, the legal gymnastics, you see some really good appellate scholars and lawyers that just come from mass training.
Not many, but like Lawrence Tribe for one, you know, Harvard Law professor, one of the, you know, strongest like, you know, appellate litigation and SCOTUS briefs like brains on the left, Democratic.
party.
And I think there's a lot of other domains such as antitrust that's incredibly flowcharty.
Contract law, sometimes also flow trarity, bankruptcy tax, more on the corporate side.
I just love litigation side.
I mean, yeah.
So actually, just because we're talking about litigation, it's not the same thing.
But there was a ERDOH problem that Axiom saw.
I don't know if it was Axiom Prover or whatever.
Is that right?
There was a controversy about it because it had...
represented that it had solved the problem when in fact the proof had been, it had discovered the proof and then just formalized it.
So actually what happened was our competitor, Harmonic, decided to publicize that they have solved unsolved problems, Erdos number 124 and 481.
And then we trusted their literature review, believing that these problems are really, truly unsolved.
And we were a really young company at the time.
We wanted to test if our system can attempt to try the problems that our competitor can.
We fully did not expect that it actually solved them.
But it turns out that we were both wrong that, in fact, the problem has been solved before.
I see.
It's not the only time that we relied on others' literature research, and, you know, we wish on it.
The other time was this paper called Dead Ends in Square-Free Walks.
You know, Professor Miller has this problem that actually turns out to have been solved.
But we, I mean, we really should have done our part.
That is, you know.
The point I'm trying to maybe elicit is not...
Not like you guys did something wrong, but rather.
You know, there's this like Japanese like advertisement of like a whole company, like hundreds and thousands of people like apologizing in the advertisement.
It's like, you know, sorry, we raised our price by like five cents.
And that's the advertisement.
I was like thinking that maybe I should just do that.
It's so embarrassing.
No, but I think that the question of provenance of information and sort of like how do you.
It goes back to the question I was asking before about like, how do I, how am I connecting the answer to the question?
Yeah, this is a great question.
I think after the Erdo Shing, we're like extremely careful.
And so we kind of like.
You know, we didn't really look at the other Erdos problems.
I believe that Harmonix still continue to claim they have solved Erdos problems.
That might not, I don't know.
You know, there's a, I think Terence Tao and a lot of other people have a database about all the Erdos problems and the status.
I think, you know, like it is really, by the way, like it's a really easy mistake to make because there are so many Erdos problems that actually have been solved, right?
And I think.
So that's kind of, indeed, I think like, you know, search and retrieval is a hard problem.
Like you don't know if that argument or an equivalent version of that.
In fact, I think the most interesting part about that entire database is there are a lot of problems that are not directly solved, but can be just a very easy extension, almost a trivial extension of another result that has been solved or sometimes not even resolved.
Sometimes I think in this dead end square free walks case, which is.
nothing to do with harmonic, complete axioms fall, that we actually didn't realize, and Professor Conor Sander-Rorgen actually pointed to us and to Professor Miller, is that it was actually from a stack mass overflow or stack overflow post.
Like a user pointed out that there is a 1936 result.
It's fascinating.
I think it's hard to find out.
Why search is a hard problem.
I guess that means that you do, does the, conjecture engine or whatever, does that use search as part of its process or is that something that you kind of, the human does and then feeds?
I think knowledge graph or knowledge base is a very, you know, important component of any company.
Yeah.
And I don't think it's talked about enough.
And so, and you guys, it sounds like you don't want to give us too many details, but like, so you guys have a knowledge graph.
I mean, that brings up also, I read somewhere that you guys have a really massive database of lean proofs that you've generated.
So synthetic data in some sense, but that this maybe is a competitive advantage for you.
I think everyone is trying to accumulate like a data, which is not a mode, it's just time and time mode.
Yeah.
It's all about like, you know, whether you can execute fast enough to make sure that you have like a certain buffer because of, say, your data set, you know, accumulation.
But that is only just a buffer.
Have you ever thought about doing something like an alpha zero for math where you start from nothing and let it just make up axioms and see what happens?
Ah, this is a wonderful question.
I think that's a very interesting approach, actually.
Yeah, I think we believe in something, which is that, like, you know, suppose action prover can be a really strong mathematician, and then really the thing that it is proving every day should hopefully help it improve, right?
I think this sort of self-improvement is extremely valuable.
And I think there are other people in the AI for Mass community.
I think Professor Gabriel Parashas' work is very interesting.
I think there are some of the kind of more conjecturing type of exploration.
Suppose we just kind of change, you know, a lot of the...
There are specific things you can do in certain ways that can try to see if your system can learn to conjecture and build theories.
I think that the topic is really...
interesting and important because it really, you're claiming that to get to super intelligence, there's sort of this like, it's just not going to be possible.
Maybe if you had infinite resources, you could just RL and it would work, maybe.
But the reality is, is that you just can't be sample efficient enough or whatever it is to do that so that you need some sort of verifier in the loop.
with the inference process rather than because you do have verifiers and like sort of during the training process and you just don't have them during the inference process.
Yeah, I think a lot of them are just secretly like trying to use this to ground their reasoning.
Yes.
I mean, I would.
I was surprised that like when O1 was...
you know, everyone knew O1 was coming, but it hadn't come out.
I was sure they're going to announce that they're using Lean to do like formal verification of proofs and actually generate proofs and then verify them so that they're grounding and reasoning.
I mean, that was my intention.
When E-mail was there, there was GPTF.
That was a great piece of work.
There's also MiniF2F.
These are all formal math work at OpenAI.
Okay.
So presumably those guys are doing something.
No, no, they all left.
Oh, they all left.
So that's my point, which is that if you're like, you know, an intern, I guess you can't be an intern forever.
So let's say you're like a junior, you know, like member of technical staff and you want to work on something for like as long as it takes to solve it.
Weirdly, people think about startup as this sort of your runway can just run out and it can just like all fall apart thing.
You might have a better chance of staying focused on the same problem for as long as it takes at, say, a startup like Axiom or one of the other new labs.
Yeah.
If you're aligned to the mission of the...
Then Big Tech.
Rather than, like, somebody decided that what you're doing is no longer...
Yeah, yeah.
It can be your VP lost some political fight.
And so, yeah.
Yeah, absolutely.
No, obviously, if we succeed, then they're all going to, you know, start doing that again.
Yes.
And then, like, I guess as a talent, then there are more, like, you know, potential places to choose from as well.
Yeah.
So then your job is real fast.
So they're struggling.
So actually, we haven't talked about it, but you actually also just released.
an API for doing lean verification.
And I actually tried it with Cloud Code because it's easier than setting up your own lean tool chain.
And, you know, like try to get lean to prove some stuff.
And the infrastructure is maybe non-trivial, especially at scale.
So you want to talk a little bit?
Yeah, yeah.
So we just released Axle, A-X-L-E, stands for Axiom Lean Engine.
And it's really a set of proof validation and manipulation tools.
that are built for Lean in the language of Lean.
So it's a bunch of metaprogramming tools.
Now, metaprogramming talents are extremely, I think, like, you know, hard to find.
And we're so grateful to have, like, a really crack team working on that.
And we want to kind of, like, release it to the community to use for free because we think that there are probably other people doing also, like, large-scale Lean operations.
And these tools are going to make their stuff go a lot more robust and faster.
So at scale.
And Excel is currently, I think, 14, like, such tools, starting from verify proof, which is the sort of to make sure that there's not nothing weird, you know, going on, like no sort of cheating by link code.
You don't axiom something out, you know, you don't assume weird things.
If you axiom N plus N equals N, you can prove two plus two equals two, which you are, you know, for sure not, that's not the right answer.
There are also, like, you know, a lot of other kind of of generation tools.
For example, you can try like different repair attempts.
So, you know, broken Lean in and then good Lean out.
And, you know, there are like currently, you know, other repair methods by IOM.
So hopefully this, what we provide can be just a lot cheaper and more kind of, you know, straightforward.
And it's just, you know, I think strong and better engineering can get you to a place that's quite far.
A lot of the people from the Lean community has been using Excel, even if it's just been a week to do all sorts of things.
of different interesting things.
We have seen people from the kind of blockchain community use it to do interesting things, et cetera.
And we have seen also, we have heard from a lot of the people that Cloud plus Axel is kind of their go-to setup for now.
We think that these are really interesting tools.
I think famously, I think today there's this mathematician who said he formalized the download news, you know, using Cloud to prove.
I think, a result, Ramsey result, and to formalize the lean proof.
And then that is also using Excel tools.
So we're really glad to see people kind of already using it.
I mean, I feel like this is a great opportunity for the collaborations that Terence Tao was talking about as well, where once people have access to the common tools, then it becomes easy to do.
And I mean, like if you have an intuition, even...
not a strong mathematician like myself.
you might be able to participate in the, you know, sort of like an effort to prove a larger theorem or something like that.
Yeah, I think that's very interesting, like, which is that, like, if you think about, like, mathematics has been not, like, as collaborative as software engineering.
You don't have, like, hundreds and thousands of people working on something together.
I think polymath was an instance when that happened that was fantastic.
So if you have a lot of really good sort of setup, indeed, like, commoditized kind of access, then people can all participate in.
In fact, that's how...
I think some of the large formalization projects have been done.
Things are divided into subtasks.
But really the blueprint writing process by, say, Terence Tao and Alex Kondorowicz of...
assigning the task to different people and how things kind of fit together, that blueprint writing part is extremely important.
And there has been, I think, a result about sphere packing, I think, by one of the other companies out there.
And the blueprint part for the A-dimension is still pretty much built on what the sphere packing community, the link community, the humans blueprint, and similar with some of their other results as well.
The blueprint part has still been human generated.
And I think auto-generated blueprint is going to be.
a technical bottleneck that many people are trying to solve around the same time.
So is there value in me as a, you know, cloud code user trying to attempt like some small lemma or whatever where I don't have a great understanding of the math?
Maybe I have a high level understanding.
Depends on what are you trying to formalize or are you trying to prove?
To prove new things.
That's a good point.
Yeah.
So maybe you would obviously probably start with formalization, right?
Yeah, yeah.
You already know the proof and you just can't get it.
Nobody has been able to get the formalization correct.
I do actually have seen people use lean and formalization and they try to do it by hand, you know, not using any AI as a way to learn mathematics.
No, it's, you know, it's all the formalization.
You don't have that.
process.
Well, it's interesting because I think a lot of my friends who started, you know, working on Lean and Math Lib was because they are in PhD and the problem is really hard.
We get stuck all the time and we want to kind of review some of the undergrad classes, a time where we still understand what the math was about and we do so by, you know, doing Lean.
And I think that's very beautiful.
The material.
Yeah.
But if you have, for example, like, you know.
access to action provers that also can formalize, auto-formalize things, and you don't have, you lose that part of the learning process.
Yeah.
Yeah.
But I do think that, you know, like for, you know, you and I, we can set up like Axle and try to like see, you know, what results we might be able to prove.
And I think that's quite interesting.
And thanks to Axle sort of making the speed a lot faster, you don't have to wait.
very long.
I remember the Pundam exam day.
We were all in the war room.
It was a Saturday.
We're all really excited and we just got the exam paper from the official organization, the proctor of the Pundam exam.
We just were looking at how much workout Axel was getting.
And without it, we couldn't have solved it with, I think, the eight problems within the time limit.
Definitely not within the time limit.
And I think one thing about these tools is like...
It's very interesting in that potentially you can have interesting reward for RL as well.
What do you mean by that?
So, for example, verify proof can be a reward for just basically a proof is completely correct and validated.
I see.
I think formal verification tooling can be interesting direction to pursue with RL.
Yeah, so you mean, for example, you should auto-formalize the informal proof and then verify and then use that as a reward?
Or do you mean?
No, as in like you pass like lean programs in these formal tools, right?
Like, and you will have some sort of score.
Okay.
Yeah.
I think if I were to build a one or something, I would have, in my mind, I would have used what I just described.
But you're saying just to learn how to do lean.
Sure, sure.
The value proposition, which is interesting about Frontier Lab, is that suppose you are a 2C business, then sure, you can just not do what we are doing.
And we have seen, for example, DeepSeek originally having a formal team and then later to solve that team because of strategic direction change.
That's all completely reasonable.
Now, suppose you are focused on coding, right?
And you have...
talent who want to work on what we're doing, it makes a lot more sense for you to do code generation, further your strength and moat.
You can partner with Axiom, just like how, for example, Frontier Labs partner with startups that work on search, such as Excel and Parallel, right?
Just call Excel API for search and potentially, you know, if you're a Frontier Lab, I think you should call Axiom API for verification.
Vital proposition.
It doesn't make sense.
I mean, it's just, you know, potentially, I think the talent, the finickiness of lean, the sort of data code, like, you know, there's no reason to.
Yeah, I mean, it took me five minutes to set up.
Why did you decide to start Axiom?
Right.
Why did I decide to?
You were a grad student at Stanford.
Yeah.
And, you know, in math.
Yeah.
So what made you decide?
I wasn't in math for very long.
I was, I think like almost as soon as I started the PhD, I just started fundraising.
So it wasn't like.
Oh, really?
Yeah.
Was that the plan or did you, did you start there and you're like almost immediately realized that.
Right, right.
So the year of law school, right, it was very, very interesting to me, like on the intellectual level.
But it's also the first year where I had no science, technology, math whatsoever in my life.
It's a weird year, right?
Like I'm reading a lot, I'm practicing, well, I'm learning how to write, I'm learning how to read.
But like, I'm just kind of...
I want to like be obsessed about something in technology.
Like that was also what's going on that year.
So yeah, the year of law school, right?
And it was very, very interesting to me because it's like, okay, like I just, I need to be obsessed with like a technical thing because otherwise I get to, I don't think I'm bored because I really love like everything about law.
I really, really loved it.
It was something that's incredibly interesting to study.
But I just...
I mean, I've been basically like, you know, very excited about like the progress of reasoning.
I was looking at a lot of the post-training kind of papers.
I was learning all of these like just by myself.
And then at one point it got to a point where I'm like, I think this is for sure happening.
And like, I think talking to Shubo, right, at birth, like every weekend, also like, it didn't help like soothing this thought.
So I got more and more obsessed.
And at a point I'm like, okay, if I'm doing this like literally every minute and I can't think about something else, like, you know, I need to do something about it.
I mean, it's like I fall mathly in love with the idea that AI is going to do math.
And like, okay, now do I do math?
It's really, really crazy.
Like at a time where I remember the obsession was quite, I just couldn't get out of it.
And then I went to this Night Hennessy event.
Night Hennessy Scholar Denning House hosts all sorts of free lunch events.
And those are great because you get free food and you get interesting intellectual exposure to things.
And I remember Julie Jor, who was, I think, a Facebook, first Facebook PM, came to speak.
And then after that, I just basically walked up to her and I said, like, what do you do if you want to do a startup and you really wanted to do academia because you kind of love math?
And then she's like, well, you know, what's your time spent on these two different things?
And I'm like, 100%, 0%.
And then she's like, well, you kind of have to follow your energy.
Yeah, I mean, if you are completely obsessed with it.
Yeah, I was completely obsessed with it.
I thought it was going to be big.
And I thought like, it just has to be a for-profit startup because like, It's so much broader than making mathematical breakthroughs.
If you think about like recursive self-improvement and like really the kind of more high level like concept of like you really want to have just AI, AI scientists.
Like the math reasoning is going to be, it's going to be a pretty.
big part of it.
And now I think the sort of belief by Cursor and Claw and other folks is like, okay, just like math transfer to coding, coding transfer to math as well.
I think that's true.
It's just that, like, you know, why not push it directly?
I don't get it.
You need to push that directly.
And then there's this other, like, you know, thought, which is that, and maybe kind of going back to the collaboration point, right?
Verification has traditionally been thought of as, okay, well, there are some industry where there's a lot of guardrails.
So if you're working in defense, military use, okay, you need to like basically satisfy a lot of barriers to entry to meet those stringent requirements.
So it's something that verification is for the industries that are closed.
But it's for the first time now, I think, verified AI is to open up collaboration.
Either it's human-AI collaboration.
Well, before blueprinting, that's human-human collaboration.
And lean was the grounding, was the verification, formal language.
And then human-AI collaboration, like we're seeing now, future AI agent-agent-agent-like collaboration.
So, like, I think verified AI is for openness.
It's not for meeting the requirements of closed industries.
And I think, just like I think verification should not be about, oh, I remember, like, you know, there's this article, like, chatbots make stuff up.
It's AI, it's a solution to, sorry, it's a math solution to hallucination.
Verification to me is not about lousiness.
Verification to me is about scaling brilliance, compounding brilliance.
It's like just kind of going back to the collaboration point.
It's about Ramanujan being a much stronger mathematician.
He was already a really strong one.
But verification helps him extend the brilliance, like both kind of like scale up and scale out.
So verification to me is not about, you know, like erasing the mistakes that allows it's about scaling brilliance.
And the third point is that like verification to me is not about like the sort of, you know, just talking about rigor.
It's actually about performance gain.
It's not just about the stringent requirements, the hurdles that you need to overcome.
It is about like...
actual verified generation is going to make it so much better.
And I think like kind of these three points, I think the last point is that a lot of the people think that you work on verification because of your distrust for technology.
Like it sounds really well to, I think the general public, including like my parents, like, oh, why we're doing verification?
Because like, you know, technology make mistakes.
It's no, we don't think verification is based on, it's because of the distrust for technology.
It's because that's what like, expected rapid exponential scale up and the deployment and the creation of technology and technological progress is what that compels and demands.
It's a very mathematical perspective, right?
Because you're saying proofs.
Our proofs drive math, right?
A lot of math is about proofs.
And math drives a lot of science and innovation in the world.
And the innovations in math drive innovation in the world.
But it doesn't need to even go through, like, in terms of, you know, the math, solve everything thing, like, obviously stands.
Like, my point is, like, transfer learning doesn't, like, transfer learning is about, like, pushing math reasoning.
It just...
So there are kind of, I guess, like there are a couple of narratives here.
Like for some people, it's like you solve math and then math are the, you know, fundamentals of science.
So that's actually the, from AI for math, like take this radical layer of AI for science, it's that narrative.
We actually believe in just like general transfer learning.
Like I think Axiom is on the infrastructure stack.
And you think that this is just a first step to, you know, basically...
unlocking capabilities in many domains in science and law, for example.
Yes, I think it's, so again, there are like, you know, multiple, multiple kind of like beliefs.
One belief is that there is math and there is like, you know, formal, the power of formal verification.
Suppose we actually, you know, solve math and have a really strong informal math reasoning engine.
We do not expect that TAM to be as large as solving math through the formal way.
Why?
I mean, code S, it is language, but it is indeed on the more structured end.
Yes.
it bridges informal and formal.
Yes.
What we are doing is it's not informal versus formal.
We're not taking this sort of like completely formal of a proof approach.
Like it's bridging between informal and formal.
It is bridging between high level and low level.
It is a direct, it's sort of like a direct improvement to reasoning through transfer learning.
And it's also indirect in that like, okay, well, like math is going to unlock a little science and sure.
And.
That is really what we're seeing.
So you think that it enables transfer learning?
Yeah.
I see.
I think that is pretty much a consensus.
I think it is a consensus and there's a bet that has been pretty much kind of overlooked by others because math sounds pure and it doesn't sound like there's any commercial value.
Well, I do obviously understand the opportunity cost if you're like a really like a frontier lab of solving this problem.
But I definitely think it's a problem that if you're like a well-resourced startup, you should be doing.
That's an interesting perspective.
Did you get everything else that you wanted to say?
Yeah, I think it's like, you know, like the question of like, is Axiom math or is Axiom verification?
The DNA of the company is math.
We think verification is the best first market.
Yeah.
And we think that sort of like solving math and especially like formal math is going to like help us like tackle the really ambitious quest of verified AI.
Now, when we are done with that, we might have other, that second market, including effort science we just talked about.
But on the theoretical layer, I think real world testing is important and potentially we can stay in the digital world and software stuff and for other things to be getting real world, like physical world signals.
But do you think that the capability of doing really powerful reasoning, once you have that powerful verified reasoning engine, that that's the moment when, okay, now we've unlocked that for...
you know, software verification and hardware or whatever.
But, okay, so now what about biology?
What about chemistry?
So that could be one.
The other one is then like really how far are you to recursive self-improvement?
Okay, so just AGI.
Yeah, I think there is this sort of question and different people because of their probably different backgrounds have different, it's really where your energy and your...
passion leads you.
Like for some people, actually, I have heard this actually, you know, with my friends, they want to work on AGI because they believe solve AGI, solve death.
There are other people who come from a more like medicine background.
They really believe they can solve death and they don't solve AGI and then solve death.
They just solve like AI for science.
Yeah.
Now, which way is correct?
I don't know.
And so the recursive self-improvement angle, it sounds to me like you're saying that The combination of verification plus the sort of like language, which is informal, it's that combination that enables really good recursive self-improvement.
I think recursive self-improvement is going to happen anyways.
We're trying to have like formal verification earn its place.
So we like, again, the whether formal verification can be welcomed and deployed and become a consensus depends on how well we execute.
And I think when you boil down that problem into an execution problem, you should just go for it.
Looking forward, what's the biggest bottleneck that you see in the field for both Axiom and maybe just the field at Broad in terms of?
Fragmentation.
So I think we're in a market where people like to start like, you know, a thousand people, they don't join force, they start a thousand things.
I think that's actually the biggest like kind of bubble indicator.
I think there are categorical bubbles and there are like other categories where there are moonshots.
It's not bubble, it just looks a little bubbly.
In the field, if people who are like really, like really legit, you know, backgrounds decide to join force and work in a team for the mission rather than for ego, for kind of the status of Neuralight founder, I think that category is really bullish and vice versa.
So I think the bottleneck actually is about...
Potentially, I think it's annoying because it's like we are in a, if you believe we are in an age of research, if you believe in like deep texts are the interesting directions to go after.
The market sort of conditions currently is good and bad.
And that good, it enables these sort of long-term, long-horizon bets to be funded.
Bad, because there's too much noise in the market and some other like irrational players.
You know, we try to work with really incredible venture firms.
Like they are the partners, they are our intellectual.
partners and there's a lot of alignment and we really bounce like very cool ideas, technical and non-technical each other like for long, long hours.
And, you know, we spend like a lot of time off work and weekend together to really intensely build the company.
But there are also other people who just want to like park like capital somewhere.
And, you know, while we don't work with them, these encourage, these are market conditions that encourage fragmentation.
And when things get fragmented, like no one gets there.
Like, I think every category, regardless of how right the idea is, is pretty much in a sort of...
earning the right to exist stage.
And if that is the case, then, for example, great deep tech company SpaceX, and people do actually join force to work on that dream and potentially in that case, also a very charismatic founder.
I think a really kind of concerning thing for me personally is that for other, probably some categories that I'm personally quite bullish about their action about and just like looking at things generally, fragmentation is a problem.
Like the sort of, you know, we see stop pulling professors from university to work on something when it really is a really interesting kind of situation.
Maybe this is a naive question, but like right now, when you're talking about players in, let's say, AI for math, where, you know, you, harmonic, and then, you know, the big labs, right?
Am I missing someone?
Is that actually fragmented, really?
I guess fragmentation, I think, is a bottleneck for the entire...
like AI landscape.
Okay, yeah.
I think AI for math is a category that is actually not a bubble because it is not fragmented.
Because people who are really amazing talents do like to join force.
So, for example, the fact to get Ken Ono and Francois Charton.
on one team.
Like, this is fantastic.
Like, you have someone who is a core contributor of Frontier Math Tier 4, really great benchmark setter, Francois, who's on the AI for Math Discovery, Proving and Discovery.
They work together.
Then you are suddenly a player with both Proving capability and Construction capability.
And that's fantastic.
And I believe, you know, as you said, like, Harmonic probably also have some really great talents, like joining force together.
I think AI for Math is a good category because of the absence of fragmentation.
But even, you know, from our perspective, sort of, for example, you know, RL, right, being, I don't think that's like a category per se, but, you know, RL talent is currently, it's quite hard to attract and retain, right, for literally everyone.
And there are a lot of companies being started and then sold like three months later.
And just each month where you could have worked on a technical problem and you're instead working on deals, it's...
a month that is wasted.
And I say that, like, you know, also with some amount of pain and suffering because of having gone through two fundraisers.
Yes, yes, yes.
Yeah.
So what's the biggest bottleneck in AI for math?
For me, for AI for math.
Not Axiom, but just the community.
But the community of AI for math.
Yeah, where is it going?
What is the thing that everyone just really wants to break?
I expect fragmentation to start to happen as Axiom and Harmonic establish category leadership.
So I expect people kind of, you know, that's one thing.
But I also think that another bottleneck could be the pressure of short-term versus long-term.
I think that we are doing things in a very sort of fast-paced manner.
But that does not mean we can always or it does not mean it is always correct to do things in the most fast-paced manner.
Like we did things in a fast-paced manner because while we were founded on the day of the International Mass Olympiad, so we couldn't have competed in that anyway.
The next Mass Olympiad is Putnam.
And we're quite excited because it's, I mean, it's an undergraduate exam.
And this year's IMO, 2025 IMO, was easy on the MOHS scale.
And PUNEM could be hard.
And in fact, it was harder than the IMO and the MOHS scale.
If you look at AI, you know, how much, how many scores the AI has retained on average and on the max, you know, difficulty of the problem.
PUNEM is harder in both, both axes.
So we want to try.
And so there's only a gap of four months, but it doesn't.
It doesn't mean I'm always going to set four months goals.
If I build a company only setting four months goals, I might build a really short-sighted company.
So there are, like, I think longer horizon problem.
I think, for example, market forces could...
force other players into chip verification.
Well, it is possible that co-verification is a holy grail.
It's possible that if you solve that, then you also naturally solve chip verification with some amount of like epsilon caveat of like distribution shift.
But I strongly believe that like a bottleneck like could be the pressure.
But I think that Axiom is fortunate that one, we are early enough to, we are like a team of just incredibly like high agency people that our execution generally surpasses expectation.
But I think like, What I think could be a bottleneck for the entire AI for math field is that potentially trying to prove commercial value is going to distract significantly from the core capability improvement.
Yeah, that makes sense.
Cool.
Thank you for driving up and coming to see us.
Thank you so much.
I know the traffic was horrible.
Yeah, thank you.
And it's been really a pleasure speaking with you.
And we look forward to seeing how things develop.
Yeah, thank you so much.
Thank you.
Awesome.
Thank you.
