# Automating Scientific Discovery with Agentic AI

**Podcast:** Latent Space: The AI Engineer Podcast
**Published:** 2026-01-28

## Transcript

MD was supposed to be the protein folding solution.
There is a great counterexample.
The counterfactual is basically a group called DESRES, D.E.
Shaw Research.
They had similar funding to DeepMind, probably more actually.
They tested the hypothesis to death that MD could fold proteins.
They built their own silicon.
They built their own clusters.
They had them taped out all themselves.
They burned into the silicon the algorithms to run MD.
They ran MD, a huge...
speeds, huge scales.
I remember David Shaw came to a conference once on MD and he flew in by helicopter and just like to this pretty famous guy, kind of rich.
And he gave an amazing presentation about the special computers and special room and outside of Times Square and like what they can do with it.
It was beautiful, amazing.
And I always thought that protein folding would be solved by them, but it would require a special machine.
Maybe the government would buy like five of these things and we could fold, you know, maybe one protein a day or two proteins a day.
And when AlphaFold came out and it's like, you can do it in Google CoLab, you know, or on a GPU or desktop, it was so mind blowing.
I forget like that protein folding was solved.
I always thought that was inevitable.
But the fact that it was solved and on like your desktop, you can do it was just completely floored, changed everything.
This is the first episode of the new AI for Science podcast on the License Space Network.
I'm Brandon.
I work on RNA therapeutics using machine learning at Atomic AI.
My name is RJ Haneke.
I'm the co-founder of Miroomics, where we build spatial transcriptomics AI models.
The point of this podcast is to bring together AI engineers and scientists or bring together the two communities.
These are two communities which have been developed independently for quite some time, but there's been some attempt to combine them.
And only now, after many years, are we starting to see some of the big developments start to...
play out in the real world and start to solve key scientific problems.
There's no one-size-fits-all solution.
You need domain expertise.
You need people on both sides of the aisle who can really talk to each other and really work together and understand both the modeling and all of the real subtleties of the system you're actually trying to work on.
We hope that we can connect these communities and that we can provide a starting point for this new era of AI and science to move forward.
So without further ado, let's get started on the first podcast.
We're really happy to have in the studio today, Andrew White, co-founder of Future House and newly formed startup Edison Scientific.
Rather than introduce him, I'll let him introduce himself.
Hey, I'm Andrew from San Francisco, former professor, now running two startups.
One that's a nonprofit research lab and one that's a for-profit venture-backed company.
And we're trying to automate science.
We're going to get into all those points.
I'm really happy to be here.
Thanks for having me on.
I want to know personally about jump from academia to industry and quasi-industry.
So I would love to hear that story.
Yes.
I guess that's the whole story, right?
So I did my PhD at University of Washington.
And I worked in a group with, I think, 19 people doing experiments and like two people doing simulations.
And I was working on a topic called molecular dynamics, which I think is actually suddenly becoming interesting again as everyone's looking for ways to generate data from first principle simulation.
And molecular dynamics, you know, covers basically everything that's molecules moving around in dynamic systems, so like biology.
Of course, the complement in material sciences.
Things like density functional theory where you can model chemical reactions in these like solids.
So I was working on that and we worked on biomaterials.
And so the goal of my PhD was trying to find what are called non-fouling materials.
So in biological systems, whenever you put like a foreign object into the body, it will trigger a response.
And that response called the foreign body response basically encapsulates it in like this layer of collagen.
This actually is exploited for some implants.
Like if you get a heart.
sorry pacemaker installed like it coats it with this collagen so that if you go to change the battery you can almost change the battery out like without even bleeding because like the body has like completely encased and this is great for pacemakers but for like a glucose sensor or like a you know brain cognitive interface bci is what they call it now yeah they're it's not so great and so that's why some of those things have like a limited lifetime because eventually your body treats it like a wound and heals rejects it yeah it's it's kind of like Some rejection is like immune-based.
And so that's where like if the body can see anything on it, like if it can see like some ligand that it can bind to with antibodies, then you get this like inflammation, which is like a rejection response you see in organ transplants.
But with materials, the body's just like, oh, there's just like a wound or there's just something here and it just covers it up.
I think, you know, the research in that field has gone on a long time since I left my PhD.
And there was a lot of theories about it's related to the mechanical properties of the material.
Like if it's spongy, there's things like if it's trabecular, like it has a bunch of little pores in it.
We worked on the theory that it had to do with how hydrophilic the material was.
But anyway, so I was the only one working on computers in this group.
I couldn't figure out like how to connect what's on the computer with what's done in the lab.
Because you can make like a simulation of whatever, 10,000 particles, 10,000 atoms.
It's like, well, this is not going to model the human body.
A lot more atoms involved.
So I had a good time.
We did some cool stuff, some bioinformatic stuff.
I learned a lot.
But then when I did my postdoc, I was like, okay, we're going to try to merge experiments and simulations.
So I worked on this theory called maximum entropy.
And it's about like, how do you take complex simulations and match them to limited observations?
And it's like the inverse of machine learning.
Machine learning is like give simple models.
You're going to a lot of data where I had like complicated models and trying to fit to very little data.
Yeah.
It was fine.
It's great.
We wrote some papers.
It was useful.
And then I wrote, I started my research group at University of Rochester on applying these methods to model peptides.
Yeah.
I'm always like too early for things.
We studied peptides for, I don't know, four or five years.
And it was...
A cool niche field, not that popular.
Now peptides are like the hottest thing ever.
I think there's even like a peptide rave I heard about a couple of weeks ago.
But when I was an assistant professor, nobody cared about peptides.
So we worked a lot on different ways to combine them.
We looked at like different experimental methods that we could do these molecular dynamic simulations of peptides.
And then in 2019, I was like out on a sabbatical at UCLA.
They have a place called the Institute for Applied Mathematics there, which is like...
This institute where people can go and do a sabbatical and learn new methods.
And they happened to be doing machine learning for physics.
I think the name of it was like some symmetric thing.
It's like machine learning for physics and physics of machine learning.
It's kind of a cool concept.
But like Jan LeCun was there and Frank Noe was there, who's a big guy in Europe in this field.
I don't know, it's just Terence Tao came by.
So the great group.
Yeah.
And everyone's kind of jamming.
It was like 2019.
So like they're not really been the big hit, especially in non-computer science fields.
Right.
And then I came back from that.
I was like, well, I got to teach class on this.
So I'm writing a book about like how you can apply these methods in chemistry.
It was very kind of niche field because every machine learning class.
that my PhD students could take at the time, this is when I was a professor at University of Rochester, it was always would end in like, okay, this is an RNN and this is like what you need to know.
Or like, this is how you do image classification.
But in chemistry, it's all about graphs, right?
It's all about how do you represent these graph structures.
It's all about symmetry and geometry.
And that was like not a thing.
It was very popular, but you had Max Welling on before.
The godfather of geometric deep learning.
So I wrote this textbook about like this, these methods, and there was a bunch of interesting mathematics to it.
I had a good time and stuff.
And then I think I was following, you know, the news in the space and codex, the original codex came out.
And I had been looking at Transformers for a while.
I just tinkered with them.
And we started trying them on doing some chemistry tasks.
We were really impressed, actually.
And we wrote a benchmark.
And this is like 2019 or something.
We wrote a benchmark of verifiable rewards in 2019.
Maybe it's 2020 by then, but like.
Ahead of the curve.
A little ahead of the curve, yeah.
Like, here's a function.
And, sorry, there's like a task, which is like, I have a body of a function for like a Markov chain Monte Carlo simulation.
It's missing some pieces.
Complete it.
And then we had like a verifier that would see, is it a valid MCMC simulation?
Yeah, yeah.
We wrote this paper.
It ended up coming out, I think, in 2021, 2022, because it took a long time to bank enough questions.
But I wrote an opinion piece about how...
transformers could change how we think about chemistry and things like this and how we teach it.
And then opening eye, some people there, Lama was there.
She saw this paper and they reached out and were like, hey, we're building this new model and we think it'd be great to red team it to see like what could happen with these models if they're applied to chemistry or biology.
And so I was a red teamer for GPT-4 and I was using it like nine months or something before release, like August.
So GPT-4 came out in March and I was using it in August.
Yeah.
And then like the React and Miracle paper came out.
I think Shen Yu, he wrote that paper and I plugged it in with GPT-4 like in the fall, you know, and I was like, wow, there's so much stuff coming out.
With React.
Yeah.
And it was really exciting.
And then, so when GPT-4 came out, I released this paper called ChemCrow.
I work with Philippe Schwaller in Switzerland on this and IBM.
So that was like React applied to chemistry.
Yeah.
And what we had is we had like, there's a cloud lab that IBM built in Switzerland.
Yeah.
So we had like GPT-4 operating the cloud lab.
And then it was like I had written a literature research agent that did like agentic rag.
Again, nobody knew what agentic rag was.
I think actually Harrison Chase had like written a blog post about some ideas there.
And so I stole some of the ideas.
Really smart guy.
And basically we applied that and we saw some really cool stuff.
It was really exciting.
And then I.
We wrote the paper.
It set off this crazy storm of like everyone had a lot of anxiety about AI progress.
Yeah.
And I ended up visiting the White House.
I guess my paper was like the only time a preprint or peer-reviewed paper was presented to the president on like their schedule for like a 30-minute block.
Wow.
And the National Security Advisor at the time, Jake, God, I was confused.
Pepper.
Pepper.
Yeah, no, sorry.
One of them is a talk show host and one of them is like the National Security Advisor.
I forget which is which.
That guy.
That guy, yeah.
He like had a presentation about our paper and they presented it because there was a big tech CEO summit at this time where they sent out Sam Altman and some other CEOs out there.
This is the Future of Chemistry is Language or a different one?
This is the ChemCrow papers.
Oh, ChemCrow, that's right.
Yeah, yeah, sorry.
I probably should name these things.
Yeah, yeah.
And so.
It was crazy.
And they had me go out there and then I met a lot of three-letter agencies I didn't really want to meet.
And I'm like, you know, there's like somebody from one three-letter agency is like, how does this change explosives?
And another three-letter agency is like, how does it change breakout time for nuclear weapons research?
I was like, guys, I'm not really sure.
But it turns out that there was not that many people, world experts on AI and science.
Right.
So what's the answer?
Yeah.
Great.
Good question.
Oh, we'll come back to that.
Okay, yeah, let's come back to that.
In the end, I, you know, had a lot of energy, a lot of excitement about this area.
So I took a sabbatical from the University of Rochester, and it was Sam Rodriguez.
And Sam had been talking to Eric Schmidt and Tom Khalil, who was also a national security counsel in the Obama administration, about how to, like...
scale up these ideas.
And so Sam had this concept of like focused research organizations, which is how do you do science, like not in academia, not in like one of these kind of near monopoly tech companies, these big labs, and wanted to try this idea out.
And I was like, hey, we should do this around agents for science or AI for science.
I love Sam.
He pushes me to come up with really...
lofty ambitions.
So we decided to automate science as the goal instead of like, see what fun stuff we could do with agents in science.
But I think that was maybe the real mission.
But of course, automating science is the long-term mission.
Yes.
And so that was what led to Future House.
And that was a very long-winded...
Yeah.
No, no, that's great.
So you chose to leave a tenure track position to do this.
I was on sabbatical, which is a beautiful concept.
But then I did resign my tenure position when we co-founded Edison.
Yeah.
Okay.
I see.
It was, you know, I had been on sabbatical for a very long period of time.
And so at a certain point, I just had to resign my tenure.
So I resigned tenure in June.
Okay.
Oh, so that's only recently.
Yeah, only recently.
And you just felt like this is the direction of my career.
Yeah, yeah.
I mean, I got tenure and I had these early career awards, like the NSF career award.
It was great.
And I think academia is really exciting.
But I just thought that right now, this kind of area, like AI over sciences, Just, I think, A, difficult to do in academia and B, so exciting that I think you can take bigger bets.
And I think having a tenured position and writing research grants is maybe not the biggest bet you can take on a field.
So now we have a venture-backed startup called Edison, which we spun out of Future House.
And we took a lot of the ideas and we're trying to do this at an even bigger scale right now.
And so Edison was always kind of the plan, like going back to Sam's idea of a FRO or like was a fundamental research organization.
Like he always had this goal of like, let's do fundamental research in this like tightly scoped nonprofit, which can kind of explore.
And then you have that as a natural arm for spinning off.
Yeah.
You know, venture backed.
Yeah, I think that's right.
I think some things that like make that.
not as clean these days is how expensive AI research is and how expensive GPUs are.
So I don't think we can repeat it many times from future house.
It might be like an N of one thing right now.
It just may, maybe not.
I don't know if venture capital keeps growing, then maybe, I mean, maybe we can, but yeah, I think we took a lot of the ideas of future house.
Another thing like, I think we expected it to be harder to automate science and actually it's really hard.
I feel like I'm always miscalibrated in this domain, but it's always hard to predict progress.
And I think that I overestimate the speed of things on month scale and I underestimate things on year scale.
So the two years from 2023 to 2025 was an enormous amount of progress.
It always felt like things were not going as fast as I thought.
But when you look back on it, wow, there's a lot of progress.
And so I think the idea that I think in Future Health Center...
Sam actually regrets us writing this, but in the original marketing or like the announcement is like, it's our 10-year mission to automate science.
And like now it's like, okay, yeah.
So two years later, we had Cosmos and things are going so much faster.
And also this kind of, this thing, you notice it in San Francisco.
It's like, it's actually kind of hard to find problems which are like so hard that like they are a challenge for language models, but not so hard that they're impossible.
We're in this gray zone, and actually I feel like that's where we are right now is that we can actually automate so much of the scientific method because it turns out, especially in a field like biology, which is very empirical limited, the top 1% guesser of what they think will happen in an experiment, the top quintile or quartile, they're about equal.
And so, you know, even if you had an even we wait 10 years and get even smarter models, I don't think it's going to really change the fact that we're ready to automate a lot of science with with existing LLMs.
I mean, what do you mean by automate science?
That's like a pretty loaded state.
There's lots of things like there's many ways of thinking about that.
So we try to draw a line between what I would call groups that are trying to model something like the cell or how proteins fall or like how antibodies can be designed or like maybe the virtual cells.
Like an example, if they're trying to use.
machine learning or AI to model some very specific system.
We're trying to automate the cognitive process of scientific discovery, making hypotheses, choosing experiments to do, analyzing the results from experiments and using it to update your hypotheses or your confidence in those hypotheses.
And then leading to a world model, which of like, okay, this is how I understand this process to be.
And then that begets new hypotheses or new experiments.
We want to automate that sort of loop.
We thought that we would have to build up like a whole new organization from the ground up for agents so it means like automated labs it means like putting all the papers in one spot like getting apis wrapped around everything but um over time like the models have gotten better and better that we had to like you know stop and rethink okay we don't actually have to hold their hands so much anymore or like they don't actually need to necessarily have an automated lab they can like write an email to a cro or something or they can like tell you what experiment to do and you can take a video of you doing it and show it to the model and they can be like okay well this is you know what happened so It's been a really interesting experience of like sometimes, you know, we over-engineer things and sometimes actually basically just mostly over-engineer.
So I always think about systems and scientific is a system, like scientific process is a system.
I always think of it in terms of constraints, right?
And like, what is a bottleneck in the system?
So what is your hypothesis about this, right?
Like in my mind, not knowing a ton, but in my mind, the constraint of the scientific process is, the work you do in the lab.
And that's sort of notably missing from, well, not entirely.
You said, you mentioned automating lab and whatever.
So like, how are you thinking about this?
Yeah, I think you're right.
Is that basically the best model, whatever, Opus 7 or GPT-10, like it really can only propose the first experiment, maybe slightly more clever, but at a certain point you just need information, right?
Like some little calculations you can do that like there's more atoms in the brain.
than you could ever simulate even if you had all the energy from the sun.
I think you could simulate maybe a thousand brains in real time with all the energy in the sun because there's just too much information.
So science really hits these bottlenecks where you just actually have to go measure things.
We definitely think about maybe lab-in-the-loop sort of situations.
One of our papers, which was called Robin, is that we had one of our agents propose an experiment.
We did the experiment, and then we had our agent analyze the experiment and propose the next experiment.
And that kind of loop, I think, is where you want to get to.
What is the bottleneck in that?
I don't think it's like the intelligence of the first experiment.
I think the bottleneck might be something like right now.
I think the bottleneck is something silly, like knowing what's the lead time on all the reagents that you need and what is available in the lab.
I think whether GPT 5.2 Codex Max or Opus 4.5 is going to do better probably doesn't matter.
It's just a matter of which one's going to have all the information about what's in the lab and how much will it cost, how long will it take.
And also, I guess, the kind of...
frontier that i think about for these models is taste which is like a lot of science i mean of course we want to you know accelerate technology we want to improve the economy we want to improve people's life expectancies want everyone to be happier but a lot of what is done in science is is based around like human preferences like why do people study i don't know a particular worm well like there is a theory that by studying the worm it has led to good medicines or it's led to discovering new genes but also People studied in the past, people's careers depend on that worm, and people want to write papers about that worm.
And so there's a human element to some of this.
And I think that models don't capture that so well about knowing what is an exciting result and what is a boring result.
I see.
So I think that's like a scientific taste.
It's like a broad category of all these things.
How do you, like, do you try to quantify taste in any way?
I mean, I know that I have some, like, fun anecdotes about this, but maybe, yeah, I'd just like to hear what you...
Yeah, actually, we sat on this idea.
We sat on it, but we, like, argued about it for a long time.
Sam and I usually, every Monday morning at 8 o'clock in the morning, Sam and I meet, and we're both, you know, caffeinated and ready, and we argue about stuff like this.
And we had a lot of Mondays where we talked about scientific taste.
And in the end, we're like, okay, let's just do the dumbest thing, which is to, like, have our agents make hypotheses and put them in front of humans and have them be like, I like this one or I like that one, right?
So we just did, like, whatever, RLHF on hypotheses.
And we learned a lot about how bad our LHF is with people.
Just like people pay really attention to the tone, to the details, to like how many specific facts or figures on the hypothesis, right?
Like actionability about like if the experiment is feasible.
But what people didn't really pay attention to is like, I don't know how to describe it, but like if this hypothesis is true, how does it change the world?
If the hypothesis is false, how does it change the world?
This like...
How much information do you gain?
It's not really information, but like impact or something.
And that really didn't come through from those things.
So then we're like, okay, well, this is maybe one strategy.
And so we had to go back and think about it more.
And we then took a pause from that research and then we made Cosmos.
And then Cosmos like has baked into it taste, right?
Like at the end of the day, there will be some report and we're working on generalizing this.
Basically at the end of the day, we're like, okay, I made these discoveries.
And a person would be like, great, I'm going to download that one.
Or I like that one, right?
I don't like this one.
rolls up to some hypothesis that came earlier in the process.
And so we think we can get to end-to-end on this as opposed to human preferences.
So you mean the feedback loop is the click?
It could be the click.
It could be also like the, you know, we do an experiment.
Sometimes in Cosmos you could ask to end an experiment and go see what's the experiment is success or failure or something like that.
But I guess like we brought it out of this kind of like hard to quantify is this a good hypothesis or a bad hypothesis and into this like you can see some downstream consequences of the hypothesis.
So, yeah, humans have, I think, a very, like, strong, well-calibrated nose for science.
Like, I mean, maybe you could argue there are sociological effects, like, across the community, but ultimately, oftentimes people really, like, good scientists know right off the bat, like, is this going to be likely to be useful or not?
How long, how many attempts did it take before you started to, like...
see results that to yourself seem useful like even working on this for i guess two years now you know i think when the co-scientist paper came out from google i think it was a really interesting idea to do this like tournament style or just pairwise ranking of hypotheses right so i think co-science is very interesting counter example to what we built is that what we built is something with either lab in the loop or data analysis in the loop or literature research in the loop where you're like iterating on an idea I think co-scientists took this like very different approach of like, let's list all the ideas and then try to come up with a filtration process to come up with the best hypotheses.
So co-scientists will produce these very long reports of like, oh, we like really tested this idea with lots of dialogue and it's very interesting stuff.
And I was really impressed with the paper that came out.
And then we had this Robin paper.
And one of the things that came out of the Robin paper is that the hypothesis that people thought was best was not the one that led to success in that paper.
Interesting.
It was in age-related macular generation or ocular age-related macular generation.
Basically, it's like part of the eyes.
You're going blind because you have this accumulation of debris in the eye and you can't clear it out.
That's one of the major cause of blindness in people over 60.
Ollie, who works on the hill.
Yeah, yeah.
He'll cringe when he hears me say that.
Something like that.
Something like that.
Sorry, Ollie.
And that one, like we went to optometrists or ophthalmologists.
I'm actually confused on that as well.
Sorry, Ali.
But essentially, you know, ask them, what hypotheses do you think are good hypotheses?
What do you think would lead to like a good mechanism for treating dry IMD?
Yeah.
And yeah, there was, you know, they agreed to be on the top 10, but beyond that, it was kind of noise.
Yeah.
And then, you know, what we found was ribosutal was a very good medicine and had a mechanism that I think is novel, although there was lots of debate on X because I think in...
2012, there was a master's thesis which proposed this mechanism on like page 38.
I actually think it was a typo.
I think they meant wet AMD.
But anyway, I won't belabor the point.
I will concede that maybe there was one reported example of it in the past.
That was a really eye-opening experience for me because that was a...
The first really serious test where we really went to the lab and we spent like four weeks on a battery of experiments to see what hypothesis led to a good mechanism and a good repurposed drug.
And it was not as correlated with human opinions as I expected.
And so since then, I think that I have a lot more faith in these like verifier in the loop.
kind of scenarios where you have either data analysis, literature search, or you're running a unit test or whatever, you're going and running the experiment.
Anything like that, I think, is going to give you a higher signal than the sort of vagaries of like, oh, this is a higher opinion or we like this one better.
Yeah, Maxwellian called it nature's computer.
Yeah.
It's like you have this computer computational cycle you're running and nature is part of that.
Yeah.
I'm curious.
So you said that there is a paper which maybe like could propose, maybe propose where this molecule came from.
But like, do you have some way of interpreting or like understanding where that hypothesis originated in the absence of that?
Is there a traceable thought train?
Yeah, yeah, yeah.
Actually, this is something we pay really close attention to.
At Future House and at Edison was provenance of like information.
So our first sort of like agent was paper QA.
Sorry about the name.
Paper QA sounds like an email set, but it was an email.
It really does.
Paper QA has like every sentence that it outputs has a citation to a page, right?
So it's like a lot of provenance.
And then we basically built along a philosophy for everything.
So Robin would, which is the name of this, I don't know, workflow or something, you can call it, that led to this result on Rippecutal being a good therapeutic for dry MD.
It has like data analysis that goes, shows you like which line, Python code led to the result here.
And then that is like, okay, then it goes to this other model, which says, well, based on this literature finding and this result from the data analysis, I believe this is the right thing.
But, you know, where does the original idea come from?
Like going after these rock inhibitors, which is the mechanism for the target was basically enumeration.
And so this is like...
If you can't be smarter, you can be, I don't know, you can try more times.
And I think that was like the theory of the Robin paper was that we can put out a whole bunch of hypotheses and then we can filter them.
Just like I think some of how co-scientists did is you go for a filtration process.
But the difference is that in co-scientists, their filtration process was other LLMs sort of ranking it with rubrics or like personas.
And our filtration process was like literature search and data analysis.
Like here's some data.
Is it consistent with the data?
Go see if anyone's discovered in the literature or if they've disproven it.
And I think that's the easy way to succeed in AI over humans is you can try more ideas faster.
Something I've heard people say, and maybe I've experienced this in my own life, like sometimes hypotheses are kind of cheap, especially in biology.
It's in many ways actually easy to come up with what you think could be happening.
It seems like to me, verifying is oftentimes a big bottleneck in maybe the biggest bottleneck.
Like if you have lots of hypotheses, you know, and it costs, you know, one one hundredth of your runway to test each one of them or something.
You don't have any shots on goal.
Yeah.
Yeah.
So how do you make sure that like you are actually enriching for good hypotheses?
Literature and data analysis, right?
You know, like, yeah.
There was a time when we used something called tiling trees.
And tiling tree is like a literal brute force method invented by Ed Boyden and Sam's PhD advisor.
And basically the idea is, okay, I want to like accomplish X.
Okay.
I could try these methods.
And then like, once you pick, I'm going to try this method, then you split into like two different paths.
I'm going to use this method or not use this method.
I'm using this method.
I need to have like, I don't know, some kind of substrate.
I'm going to try this substrate or this substrate or this substrate.
Right.
And you can basically try to really like tile.
space of all it is we tried some early experiments there and you're right you run into this thing where some of the hypotheses that come out just don't make any sense and like you are going to waste a ton of effort if you actually test them all nowadays i actually would argue that if you go to an llm and you ask it to evaluate you know hypotheses including some garbage ones it will probably do as good of a job as an expert in the field and filtering them out That's not always the case.
Yeah, I've actually seen that myself.
Yeah, but there's a lot of gotchas and I think people can miss those.
But I think they're actually pretty good.
And so I'm not as worried about hypotheses that can fail fast by an expert looking at them.
I think now the filtration process really happens in literature.
And I think the filtration process happens in looking at like, you know, like a biobank data or like, you know, what do we know from GWAS or something?
You know, other sources of existing data as much as you can draw upon.
So with regards to existing data, another contrarian take is that oftentimes the hardest part is just understanding the context of data and where it comes from and how do you interpret it.
I can also think from my own life, multiple cases where...
the data in some sense was like there and you had two people who were both experts and very smart people who looked at it and drew very different interpretations.
And in fact, like when we were interviewing Heather Kulik, she had some fun stories about using LLMs and she would find that there would be raw data in a paper, which wouldn't agree with the conclusions of the actual paper.
And it's straight from the paper.
It's not even like cross paper talk or something.
I'm going to be a really boring interviewer and be like, yes, you're right.
You know, like this is a hard question.
I think, you know, to give you something concrete, we have a bioinformatics benchmark we call Bixbench.
Bixbench, like we put it out, we've updated a few times.
It's in some frontier alums when they release their system card, they'll mention Bixbench, like one of the things they test on.
And, you know, we're getting to 60%, 70% correctness on Bixbench.
And we found that actually we're at the point where humans disagree at this level.
Like humans only agree 70% of the analysis.
And so it's true that like when it comes to analyzing data, like humans do not agree 100% of the time.
There is a certain amount of like choice that goes into it.
And we try to, so Edison is a for-profit company.
We like trying to sell some of this stuff to the companies.
And we'll go to some companies like, oh, we never impute data.
imputing data is bad, like, or, you know, whatever.
And they're like, okay, well, we'll have to change our ages so we don't impute data for them.
But then some other companies are like, oh yeah, we impute data, it makes everything easier, right?
Or, and...
You know, you want to know what the real modern dark arts are?
The, like, AI-resistant area of the world is, like, medicinal chemistry.
That is, like, the spot where, like, there's so much superstition.
Oh, yeah, yeah.
Everyone, yeah.
Everyone is, like, pseudo-religious.
Yeah, exactly.
You have to be the survivor, I feel.
Otherwise, you get burnt out.
But the religions never agree, too.
Two medicinal chemists will have completely different viewpoints about, like, a functional group.
Yes, exactly.
And I remember this as I was talking to somebody who works at a CRO, and they're like, oh, whenever, like, company X orders anything, we never put boron.
on any of the compounds because they hate boron because there was one program that was killed because there was a boron, you know, somewhere in the core and it led to some toxic side effects.
So no boron for this company.
This company, they like love things to be fluorinated or something because they love, think it's great for the Admet properties, right?
And so there's like all this stuff where you reach the point where you're at, I don't know, human bias level or human disagreement level.
And I think we're getting to that point in data analysis.
And so, of course, you will see then that if I take the raw data from a paper and I analyze it myself, I will get a different conclusion.
One of the cool tricks you can do is back to this brute force thing is that I can go to our agent and I can run it 100 times and I can take the consensus-like analysis.
Or I can say, even if you make these three different choices in your data analysis, you get the same conclusion, right?
Or this conclusion is somehow sensitive to those choices.
And then you can, there's even like words like epistemic versus aleatoric uncertainty, right?
It's like, this is aleatoric, which means like, I think it's noise from the data or this is epistemic uncertainty, which means like, I think there's some choices that are being made.
There's some model differences that lead to the disagreement.
Anyway, there's like, there's like a Donald Rumsfeld formulation of this as well.
No, no, no.
It was an aleatoric epistemic debate there.
Interesting.
This kind of.
digging into your cosmos.
Yeah.
So I, I glanced at the paper and one of the things that jumps out is that there were certain class of problems for which, uh, it was only 50 some percent accurate.
Oh yeah.
Yeah.
And can you talk about a little bit about that and how that like, okay, so if I'm just raw getting 50% accurate answers and then I'm going into the wet lab and being like, okay, try this.
And then it's like, Like the stupid thing told me to do a dumb thing.
Well, I would say, first of all, that 50% is actually pretty good because it's rare that experiments in the lab are actually coin tosses, right?
They're usually a lot more outcomes than binary.
Yeah, sure.
Okay, yeah.
But that particular number was human agreement in the interpretation.
Okay.
And so we asked people to evaluate different aspects of Cosmos.
We had them evaluate like the data analysis decisions.
We had people ask to evaluate the literature.
Like, do you agree with its finding literature?
That number that was 50%, that came from Cosmos's interpretation of some of the analysis.
Yeah.
So like, it might go in literature and find this result.
And then it would say, wow, this is super exciting.
This is amazing.
Or it might do data analysis, but this is a novel discovery.
We're really excited about it.
And then people would disagree.
That's actually not interesting.
Or like, I don't agree with the interpretation of it.
So it's like picking bad problems maybe.
Yeah.
In the negative class.
And so I think it's like that 52 or 55, whatever it is, that's interpretation.
And so I agree.
I think that's where, like I was saying, I think the frontier right now is scientific taste.
Yeah.
And so that's what we're working on right now is how do you get that interpretation to match.
Could you just step back and just introduce Cosmos from a high level?
Yeah, yeah.
I would actually be even curious to hear starting from like ChemCrow and, you know, you have PaperQA, Avery, Ether.
Zero, yeah, yeah.
I'd like to hear a little bit of the lineage and how those different decisions were made.
What were the key learnings and how did you get to where you are now?
Yeah, so.
I could retcon and tell a really great story about how we arrived at Cosmos.
But I will say that like to a large extent, we just try a lot of stuff and sometimes it works and sometimes it doesn't.
Okay.
You know, I'll say that we're very, I'm a builder.
Like I like to like build things piece by piece.
I'm probably some fancy word for it, but I'm like a Lego guy or something.
My vision was that we would make an agent that does this part of the scientific process, an agent that does this part of the scientific process, whatever.
And so we had like ChemCrow, which is going to help us with.
setting up our medicinal chemistry work.
We had ProteinCrow, which we haven't released.
I don't know if we will ever release, but ProteinCrow was like designing proteins we might need for some part of our workflows.
Or we had a data analysis agent.
Is that LLM or that's a...
It's an agent, so an LLM plus tools.
Okay.
Or we had Ether0.
It was like, okay, we noticed that the frontier models can't work with molecules very well.
So let's make a model with intuition for medicinal chemistry.
And that was what led to Ether0.
But then Sam actually really pushed on us to like, let's just see if we do the whole thing.
Let's just try to build an AI scientist.
Let's just try the whole thing.
And that was what led to Robin.
And Robin was like, let's just take these agents we already have and we'll just put them in like a workflow.
Basically, it's like you could express it in a concise Python file of like, you know, try a whole bunch of ideas, then go see if they all filter through literature or if they've been disproven and then go like come up with experiments that you could do in a wet lab.
And this is our inventory list.
And then go analyze all the data, then go back and repeat the process, right?
So that's like what Robin was.
And we came across Cosmos.
We're trying to understand what is the process that Robin is automating.
And it came from this idea of a world model, which is that when we first started Edison, we were thinking, what do we want to change about this?
What is new here?
And so we spent some time thinking about, well, the scientific process, like what is actually going on in like my brain, which is that I have some understanding of the world or the phenomena I'm studying.
And that's my world model.
And then a lot of the actions I take are about trying to update that world model.
And it's something that changes over time.
And so this is like this ability to change over time, but it's also something that is practical.
Like I can use it to make predictions about, I know from this experiment, this will happen.
That's why it's like a model and not just like, you know.
memory or a bunch of papers or something like that.
It's supposed to operate.
In Cosmos, we tried this idea out and actually Ludo, who's the first author on paper, we tried a whole bunch of ideas around world models.
And we kind of thought they weren't really appropriate.
Like, well, we tried a lot of different ways to do this.
We tried method A, method B, method C, and they're okay.
And so we all just had to take a break.
Ludo, his project...
didn't work on trying to do this world model stuff.
He's like, I'm going to keep trying it.
Ludo's a very stubborn person.
So he tried it for like, I don't know, a week or two weeks.
Then he was kind of like quietly.
He's like, Hey, can you guys come take a look at this?
And we're like, wow, this is actually really cool.
And then we like started building on it and jamming really.
And I think what Ludo figured out is that you have to get this like experiment loop thing.
You have to be able to let it in the data analysis agent is what got us in the loop.
So if you put that in the loop of like, it can really update his world model because the, we were trying to build it around literature before.
And when you build it around literature, there's just like not really experiments you can do and then see the results for.
That was like our surrogate was literature.
It just wasn't working.
Data analysis actually really lets you explore ideas.
And so that was what led to Cosmos.
And so in Cosmos, we basically, we had all the pieces sitting around.
We were working on world models.
We were working on a data analysis agent, working on a literature agent.
And then we were working on, you know, we built a platform for scientific agents.
So we had things that can write a LaTeX report.
We had things that can make nice plots.
Then we put that all together and like a world model was like sort of the...
the glue that allowed it to fit together.
An analogy is like encoding agents.
GitHub is sort of the glue.
There's some shared repo and everyone works on the repo and software engineers have spent lots of brain cycles thinking about what's the way to coordinate.
and organize working on code together for a long time.
So the world model is actually like a memory system?
Yeah, you can think of it as a memory system.
We think about it as a model.
So like it actually, you can put in input and it will output predictions.
We think about calibration.
But like really it is a set of like a big bundle of information that we accumulate over time that's distilled in some way.
And that is like what allows us to do this.
And I think you can think about like a GitHub repo is like...
It's a distillation, right?
Like really, there's a long graph of commits that lead up to it.
And like the current file system in that GitHub repo, or I keep saying GitHub, I'm such a corporate shill here.
Your Git repo is like a distillation of all of the work that people put into the PRs, into the commits.
And so I think there's a nice analogy between a Git repo and what a world model is.
I think that's just sort of what allows us to automate scientific discovery so well.
Can you talk about like kind of how you implement a world model or is that sort of like secret sauce?
That's our like secret sauce right now.
No, it's fine.
So one thing that's notably missing is the like simulation, right?
Nuclear dynamics or like bolts or.
Yeah.
I want to help you guys pump up your views here.
So I think molecular dynamics is overrated.
And DFT is overrated.
In fact, DFT may be even more overrated than molecular dynamics.
You mean for materials or for biology or for both?
For materials.
Okay.
And I can explain more about that.
Basically...
MD and DFT have consumed an enormous number of PhDs and scientific careers at the altar of, you know, the beauty of the simulation.
Also, random interjection.
Once I did an estimate, I think pre like chat GPT, something like 20% of the world's computing power just went to simulating water.
Oh my fucking God, water.
Yeah, yeah.
I had to deal with so many water simulations.
I did DFT simulations of water and they are so annoying.
I use these big computers.
from the department of defense and we i spent like i don't know five months and by the way this is pre-llm training days five months of compute is actually a really long time i simulated water with quantum you know effects with a grotus mechanism for how a proton hops through water.
And it's on YouTube.
It's my number one YouTube video.
And it represents like...
Until now.
And it represents like, I don't know, a million CPU hours of compute.
It was, you know, one of the biggest computes that I...
Probably the biggest one I've done in my life so far.
Maybe Ether Zero is bigger, but it took a lot more work.
Anyway, and what's the point?
What'd you learn?
All I learned was like, what set of hyperparameters reproduce some physical effects of water?
But none of it was de novo, right?
And this is the...
This is the issue with molecular dynamics and DFT is that they don't model the world correctly.
And so we have to invent little stories we tell ourselves about we're like making good inductive biases and then it models the world more correctly.
Like in DFT, you simulate water at 330 Kelvin when you want room temperature water.
Is room temperature 330 Kelvin?
No, it's not.
That's a little too hot, right?
And so the issue is that people just make up these...
these things are like, I don't know, GGA or like B-Lip or B-3-Lip, all these different like methods, people, but they're clearly empirical.
And then they bolt it onto DFT and they say, look, it's a first principles method, right?
But actually you made a whole bunch of choices and you know, you, you, whatever, overfit to the validation data to get this to work.
And that's, I think MD and DFT are like that because if you go look at the catalysts, you know, what catalysts changed the world, none of them are.
single crystal materials that are really well suited for DFT.
They're always like, they have grain boundaries, they have dopants, they're complicated, right?
And you never capture DFT.
So I think this is one of the fundamental, I don't know, dichotomies of the world is that simulations stimulate really boring things really well.
They don't simulate interesting things very well.
And so that's why I don't do DFT and MD anymore.
What about somewhere like the machine learning stuff like AlphaFold and...
AlphaFold was trained on X-ray crystallography data.
And I think, you know, this is the story of MD is that MD was supposed to be the protein folding solution.
There is a great counterexample.
There's a, I don't know what there's a word, but the counterfactual is basically a group called DESRES, D.E.
Shaw Research.
They had, you know, similar funding to DeepMind, probably more actually.
tested the hypothesis to death that MD could fold proteins.
They built their own silicon.
They built their own clusters.
They had them taped out all themselves.
They burned into the silicon the algorithms to run MD.
They ran MD at huge speeds, huge scales.
I remember David Shaw came to a conference once on MD and he flew in by helicopter and was like this pretty famous guy, kind of rich.
And he gave...
an amazing presentation about the special computers and special room and outside of times square and like what they can do with it is beautiful amazing and i always thought that protein folding will be solved by them but it would require a special machine maybe the government would buy like five of these things and we could fold you know maybe one protein a day or two proteins a day And when AlphaFold came out and it's like, you can do it in Google CoLab, you know, or on a GPU or desktop, it was so mind blowing.
I forget like that protein folding was solved.
I always thought that was inevitable.
But the fact that it was solved and on like your desktop, you can do it was just completely floored, changed everything.
Like the bitter lesson on steroids.
Yeah, I don't even know what it is.
But it's like, imagine chat GPT came out, but instead it was like, oh, you can just run it on your phone or locally on your own desktop.
Like that's the level of like shock that came out.
And it gets down to this thing that humans are really bad at estimating problems that aren't human-made problems.
Protein folding, we all thought, would require a huge amount of compute, a very challenging problem, the most hardest problem in the world, right?
And it turns out that you can actually do it on...
I think the numbers are now like 10,000 GPU hours.
You can train a good protein folding model.
It's actually turned out to be barely an inconvenience.
Therefore, why not?
Oh, therefore, protein folding was highly efficient based on experimental data.
They took...
X-ray crystallography, that's what DeepMind did, is they took X-ray crystallography data.
Desiree has tried the first principles method.
Yeah, yeah.
And it's like a nice head-to-head comparison.
Yeah, yeah.
Two very well-resourced groups.
They both tried different ideas, and the machine learning on experimental data beat out first principle simulation by, you know, a very large margin.
And so why isn't...
like bolts or whatever inside of cosmos like why isn't there a tool that's that can run oh we have bolts inside of we have bolts gen yeah yeah we have that inside of cosmos okay it is i mean i think in the version that we have uh for people to just sign up and use it's not in there yeah but like uh you know you can imagine that you can just modal or lambda or tamarind or 310 there's all these companies that basically wrap a lot of these these like um uh deep learning protein design tools or chemistry design tools they wrap them in an api you just get that to Give it to Cloud Code if you want.
You can give it to Cosmos and you can be like, hey, if you want to design a protein for X, use these tools.
Your mechanism, it sounds like, or one of the primary mechanisms that has been successful is it enumerated a whole bunch of possibilities and filter.
And so how do you think about serendipity and out-of-distribution thinking and getting there and how far have you gotten and what's left?
Yeah, that's a great question.
I guess the short answer is that there's very...
So this is the domain of Seaborn, so chemical, biological, radiological, nuclear weapons, or I don't know, safety.
Yeah.
This domain has been explored a lot in history by a lot of organizations.
Yeah.
And I would say that there was a big question mark for us a few years ago was like, how much of this stuff is intellectually bottlenecked?
Yeah.
Like how often are people like, oh, wow, I want to cause harm, but I need to know like some facts.
Could LLMs make that easier or go faster or anything like that?
I think, you know, the first set of answers in 2023, I think was basically no, is that like, you know, you can go find the synthesis route for many dangerous compounds on Wikipedia.
People know what are the targets in the human body that like are targeted by most biological weapons.
It's not really that much of a mystery.
So I don't think there was a lot of like, there's a lot of new ground when LLMs first came about.
Then there's a lot of concern about laboratory protocols.
Could agents or LLMs reveal some tacit knowledge that maybe people couldn't find on Wikipedia?
Or maybe for making something, there's some technique that is acquired when you scale it up in size or something.
Or maybe there's some way to get around tracking lists by ordering different compounds.
And that, I think, was really well tested.
by a few different labs.
Not me, but there were some groups that spun up that started making tests for this and labs pay attention to it.
I think it's really been put into process where LLMs will shut down or be filtered in those scenarios.
But I think that is actually an area where there is some risk.
And so I think that's something that people pay attention to for open source models.
And there's still, I think, some discussion there.
But I think to a large extent, it's not really been greatly accelerating in practice, or at least I haven't seen much evidence of it.
And again, I think it comes down to the fact that it's not really available, but if you look hard enough, you can find most of the information you would need to get up to no good in the public domain already.
But then I think now is the next frontier is like, can it somehow help you with real-time protocols, troubleshooting, like more in the loop and more, especially in the computational side of things?
There are some scenarios that are now coming into focus that could be more, dangerous or more intellectually bottlenecked and so i think people are trying to pay attention to that to some extent there was like a first wave that we thought this could unlock a lot of stuff and i don't think it came to pass yeah i think there's now an emerging sort of second wave of like there are some actually new scenarios that were just too far-fetched to consider two years ago that i think are now realistic um some smart people are paying attention to it but i don't think it's solved yet i don't know It's very vague.
So I guess like one kind of differentiator, there's a lot of talk about AI safety in like the broader LLM, you know, ASI space.
And, you know, there it's jokes about paper or paper click maxing robots or something.
But like the core threat here is more like a malicious actor using this as a tool to accelerate something dangerous.
And like kind of the first order hypothesis is that you basically already have to be an expert to effectively.
create a bioweapon or a chemical weapon and a non-expert or an expert already know how to do this yeah i think you know so so each of the categories in the cbrn they're all a little different but i think to a large extent it's a lot of like pushing material around you know the classical example nuclear is like it's a lot of a lot of centrifugation a lot of ultra centrifugation a lot of high pressure high rpms yeah and so it it's just you can maybe get smarter about how to set up the economy of scale to do that with an LLM.
But to a large extent, I think you can call your friend and country X and they can tell you what are the steps.
I don't think it's that much of a secret.
It's just a lot of moving material around.
And I don't think it's meaningfully accelerated.
Now, with that said, there are all kinds of dumb dual use things of like, maybe you want to call a company that makes centrifuges and you want to make sure that they sell you them and they go through some KYC steps.
And maybe an LLM can get you through the KYC faster.
And that's like a dumb thing that like, okay, like, yes, like, you know, email makes it so that you can order centrifuges off the internet more easily.
Is email like a dual use technology?
Like, yeah, to some extent it is.
So I think there's a lot of like weird second order things that we don't pay attention to in AI safety.
of like, does it make KYC easier?
Does it make it easier for people to know like where to order this from?
Or like, what is the expected price?
Or like, what should you order first, right?
All those like sort of simple logistical things, I think are accelerated by AI just as like a consequence of AI being an accelerating technology.
But certainly, I mean, shit guys, there's some scary stuff and I try not to think about it too much.
Yeah.
I don't know.
I guess I don't want to get too political, but I do think that right now the United States, Government is maybe taking a slower, less intensive look at safety.
But there's definitely people, I think, in other spaces than the U.S.
government thinking about it hard.
Do you think it's a thing people need to spend more time on?
I do get waves of angst about AI.
I'm sure many people living in San Francisco do get a little bit of waves of it.
And sometimes I think that there isn't enough work being done on it.
And then sometimes I think, wow, like I need to mellow out and like, you know, we have lots of time to think about it.
What is my opinion on it then?
I don't know.
I think my opinion is not formed fully.
Yeah.
You and Sam have done a lot of thinking about funding science and future of science.
You've been vocal about the reproducibility crisis and other things.
First question, why this focused research organization or FRO?
Yeah.
What does that get you that you don't get from academia or, you know, big lab or whatever?
A nice network of people.
And I think Edison is like a real, of course, I think Edison's going to be great, but I think it's a mystery of what's going to happen.
So I don't think we've had as much friction there as you might expect.
But yeah, this is all stuff that Sam and I think about all the time.
It's like, how do you balance stuff like this?
How do you balance the economics?
There are some venture-backed companies that are having cash salaries over a million dollars.
And it's insane to me that you would use all of your cash from your equity financing.
In these insane salaries.
In terms of total spend on GPUs, it can still be a small fraction of your burn.
So sometimes it kind of makes sense.
Yeah.
Yeah.
That's one way to think about it.
So this is a good lead-in to you are automating science in some capacity.
So where does that leave scientists?
So I think this is Jevin's paradox we can try here.
So let me start with a contrast here is that, you know, if we automate, you know, taxi cab drivers, there's not going to be an increase in people needing to go places.
Maybe there'll be somewhat an increase, but like there is a finite amount of like time people will be spending in cars.
And so there's an upper limit.
So when you automate that, that's like a scarcity thing.
It's basically you're displacing jobs when you automate driving.
In science, I don't think there is a finite...
appetite or a finite capacity for science i don't think science is like a scarcity thing like there's you know 100 more discoveries left to be made and then we'll be done and so like we're displacing jobs i think instead actually if we can you know make science go much much faster there will be no there will be no decrease in demand there will be actually i think an increase in demand that will match whatever automation amount we have and so my vision for what a scientist would be in the future is that they will be i don't know like uh agent wranglers or cosmos wranglers of like, okay, they're exploring a hundred ideas simultaneously, or they're like working with systems like ours to make 10 X the discoveries, a hundred X discoveries, because I think there's an unlimited amount of scientific discoveries to be made.
And so there's no like scarcity set where basically we will displace them all.
Now that's kind of like, this is what I would tell.
When I talk to a first-year PhD student, everything's going to be just fine.
But then when it gets into the nuts and bolts, I do agree that this is going to be a really hard thing where if I am CEO of a company that makes science, like a pharma company or a material science company or something like that, or an R&D arm at IBM, I think, well, I could spend a million more dollars on compute for the AI scientists or could hire 10 more people.
I might just choose to go with the AI scientists because...
To a large extent, like hiring people is hard, right?
And hiring an AI scientist is probably a little bit easier.
And so I think that there could be some friction.
But another thing is like science is in some ways closer to art in the sense that like there is a large number of people who just appreciate good science.
Like if you get published in Nature, it's not because it's really going to be world changing.
Of course, that's part of it.
But it's also because like people are like, wow, this is really interesting science.
So I think the enjoyers of science are also scientists.
And so I think that it's kind of hard to imagine a scenario when there's not scientists as the consumers of science.
And so I think if they're going to be consumers of science, they're also going to be some of the producers are involved in the process by itself.
I don't know if that makes any sense.
Yeah, you've touched on this.
The question in my mind is just what does a scientist do then?
There's a great short story by Ted Chiang.
I think in like 2003 or something.
And it's about like, well, at first scientists were displaced and they became like the interpreters of like what the AI scientists are doing.
Like the scientists read the AI scientists like papers and then translate them for whatever popular science or something.
And then after that, like they couldn't read the papers anymore.
And so they were left behind.
And so they had nothing to do and they just sat around.
But the problem is that science is like, you know, you have to translate science to make...
any impact like science cannot exist by itself i do agree there's like engineering can exist by itself like if you give some kind of system a goal of like making me a material that i can make a space elevator out of you could be not participating in the beginning the process of the middle of the process and you just come by the end and be like okay follow this recipe like science of like what's the origin of life or like is there water on other planets or you know um why is some catalyst better than another catalyst that has to be hitting human eyes and human brains at some point.
So I think a human has to be involved in the process.
Don't want to be contrarian, but...
Yeah, be contrarian.
Why does a human have to be involved?
Why does a human have to be involved?
Well, a human has to be involved at least some point to be like, yes, this is good science or this is bad science.
Okay, so it goes back to taste.
Yeah, but I don't know.
Maybe you're right.
Maybe there is no point for humans.
Maybe we'll be like, you know, what is it?
Sora, you know, like the AI slop app.
But I think in Sora, there's still humans at the end clicking on the videos or something.
Yeah, so...
So the sort of analogy kind of brings up an interesting point.
Like, is it possible that like due to the biases of AI science, if we really go full in science that, you know, there still is a market for kind of boutique human science, like, you know, there's still people who want to, you know, paint things the old fashioned way, but more to the point, does it become even more important for, to have a human who is.
actively doing their own exploration because there will be like large blind spots and biases due to the bottles that just you'll never be able to overcome because this is sort of baked in now due to your training data.
And without a human that you'll always get stuck and there will be a blind spot that will never.
Bio, which is a company in Oakland or in Emeryville, they do really cool stuff with automation.
I think they're going to be testing this theory of like, OK, maybe if that's the bottleneck.
we can see evidence of it because they're going to start doing really well.
It could be true.
I still want to say all of those in my mind are still sort of scoped in terms of like R&D for pharma or bio, but they're not like, none of them are attempting to answer big fundamental questions.
And maybe there's like different levels when I think about that.
And you seem to be, it seems like the future, how the focus of future house in Edison is much more towards like.
you know, sort of R&D and sort of end run science.
But, you know, I have some background in, you know, fundamental physics.
Yeah.
You know, it's like, is there any thought about like, how do you like take on, you know, dark matter candidates?
And like, I just, you know, think the data to really give us a complete story is just not there yet.
You know what, like, I'm sure everybody at every company, like is the biggest critic of their own product.
Yeah.
So we think Cosmos is, we think it's great, but there's a very large amount of area for improvement.
So with Cosmos can, so there's like an open, like sort of access to everybody version.
Yeah.
Do you provide access to other labs that is less open?
We have a version of Cosmos that has like bigger resources.
Like it can run for longer.
It uses GPUs.
So like basically when it does data analysis, it'll have a GPU.
We use that for things like machine learning experiments.
If you want to know this question about whether it's better to pre-train first on noisy data or not.
We have pre-release models that are coming out and we try those.
So I guess, yes, we do.
And we do have research partnerships with companies where we build something specific for them.
And that is something we think about.
Broadly, I would say Cosmos that's on the website is pretty close to what is the best we have internally.
I have a question.
So you previously have stated that you think that language is the natural language.
Language of chemistry.
The future of chemistry is language.
Yeah, yeah, yeah.
Okay, so I wonder, do you still believe that?
Good question.
I think...
I would say yes.
I still believe that.
So in that article, that opinion article, my point was that at the time when I wrote that article, which I think maybe three years ago now or something, maybe 2023, it was that we have models for predicting solubility of compounds.
We have data about large populations and we have papers and we have code.
And the only way to bridge all that information is natural language.
And the argument was that like humans, like, you know, whenever we can't bridge information, like if I can't talk about my code or I can't talk about some idea to you, I will invent words until I can get the point across.
Right.
And that humans are always innovating on language to make it represent all known observations and people innovate on language to represent whatever code pattern they have.
Right.
Like this is like the only shared activity we've been doing for this long is like coming up with words to represent everything we know.
And so I think that for that reason, natural language is the only possible way to connect all the different pieces of data we need in biology, medicine, or any domain for that matter.
I think there's some caveats to this of like, you know, you can make an argument.
Like if Jan LeCun were here and he would make an argument about like, you know, world models or like vision or embodiedness, right?
Like there's arguments against natural language that like, you know, that maybe there's something more that it's not the complete story or maybe natural language imposes limitations.
cannot exceed because you're stuck in this abstract space that was invented by humans and you can't escape it until you can like touch something.
Yeah.
I mean, it is an abstraction, right?
And like scientists basically work exclusively in abstractions to some degree.
I just, I find, I found that interesting because it seems like most scientists, you're right.
Like when they explain things, they explain things through language, but.
So many conversations, maybe most at some point, result in people drawing diagrams or something.
Like, you know, chemistry, like biochemistry largely, or medicinal chemistry is oftentimes a, it's a language of graphs, right?
Or, you know, I mean, bonds are abstractions, yes, but like, they're pretty good abstractions for most, for many cases.
Or like, you know, geometry, you know, thinking about, you know, protein is like the geometry of a protein.
It's like, I think that that's how people will, a lot of scientists like to think about things.
And so I find it interesting that like, yeah, that you are focusing primarily as language.
Like, have you thought about essentially a multimodal version of this?
Like where, you know, when it comes along a smile string, it doesn't just say, oh, this is a smile string, but like, this is a graph.
This is a representation of some higher, like abstract object.
You're absolutely right.
And the problem with these.
this like, I don't know, Jacob's ladder or something, whatever you want to call it.
It's like, yes, you can say that a molecule, you can call a molecule by its name.
You can show the graph.
Then if you go to a molecule like ferrocene, well, it doesn't really have bonds, but like part of it.
And so then you're like, well, we need to draw it visually.
And then you go to a molecule like, I don't know, glycine betaine on this dihedral angle, right?
And so like, it's not actually this thing I drew.
It's actually an ensemble between this thing and this thing, right?
Then you go to benzene, you're like, well, Not only is it like an ensemble of these different conformers, it actually has electron density, and you can't really ignore the electron density in benzene.
You need to treat it correctly.
And it's like, well, you can't actually represent the electron density that way.
You actually have to look at the correlation of the electrons individually, right?
Because you can't really model benzene with like DFT, right?
Or functional.
You have to actually look at the electron correlation.
And it's the electron correlation.
Like, well, you know, you can model electron correlation, but, you know, actually...
these things, when they're in a solution, they have like, you know, relativistic effects because it's like, there's a whole bunch of stuff around it.
So you really got to have the relativity in there.
And you're like, well, you got the relativity and you have the electron correlation.
You can have the bonds and you have the conformers, but you really think about the cosmic radiation background because like, you know, it does actually impact everything.
And there is some, some energy there.
Right.
And before you know it, you've ran out of, you know, you've ran out of compute or whatever resource you're using to model this.
And so I think, um, you have to draw the line somewhere.
Natural language, like I said, is that humans have worked for a long time to make it be the, you know, what's the word?
Like the least abstract or the, you know, it's somewhere on the border of like, it's still abstract enough that you don't need to know all these details, but it's still granular enough or concretized enough that you actually can make use of it.
There may be some other representation like multimodal might turn out the video or maybe, I don't know, there's some other like fusion that you can make.
I like natural language because we all work really hard to make it right at that boundary.
And I do agree sometimes, sometimes ideas slip and they can't be in language.
You have to get out the whiteboard or I just slip and you have to wave your hands around, you know, or maybe then, then you need that, that degree of freedom to communicate.
Just digging in on this a little bit more like famously quantum mechanics is like undescribable, right?
Like there's, there's an argument that you cannot understand.
quantum mechanics with words or with our preconceived understanding of the physical world because it doesn't behave like the macroscopic world.
And so the only way to understand it is through mathematics, right?
And I largely see language as the joint key of science as well.
But I wonder if that's not true for many domains.
And quantum mechanics is just the one that hits you in the face.
I mean, I don't know, actually.
I think the...
There's like seven principles of quantum mechanics or five or something like this that you can actually express pretty concisely in language.
I agree that like you need to actually look at the consequences of them.
You need some mathematics.
I don't know.
I actually I don't know.
This is like a challenge.
I think you could actually describe a lot of quantum mechanics of language.
Sure, sure.
But but I see your point.
And yeah, I guess I'm a realist like I.
When I talk to my kids, maybe I will be like, okay, let me draw for you.
I don't make sure in our house everything is described with natural language.
So I agree with you there.
I think maybe we can be a little flexible with natural language and include equations and smiles, strings in it.
And I think we can get a little bit farther.
So maybe that's okay.
But some people, I think, like optionality.
You know, like, oh, it could be this or it could be that.
I'm somebody I like to take strong opinions and see.
how much farther they can get me.
And I think in my career, it's actually been better for me to take strong opinions, which in my deepest of hearts, I know that are maybe not correct or not fully correct.
But once you take these strong opinions, it just, you can sort of move many steps down the road once you take these strong opinions.
And like, for example, at Future House, we took the opinion that scientific agents are the future.
And that allows us to skip a lot of steps because a lot of other people were like, we need to build a foundation model for X.
And we just skipped all that, right?
And I think if you also were unopinionated and you had optionality, like I can think of a famous example of a different company that liked the optionality and they wasted a lot of time on generation models or something, then I think you get stuck.
So that's one of my strong opinions is that natural language is a way to join all these different domains.
It may not be a correct opinion.
It may be more subtle or more complicated, but it's allowed me to get very far.
I'll drop it someday and maybe find a new one.
Yeah.
Not yet, though.
That's my meta opinion on the matter.
The Ether Zero story on your blog, I find hilarious and kind of awesome.
Yeah.
You know, when I was a kid, I loved the, like, genie slash monkey paw, like, concept of be careful what you wish for, because you just might get it.
Yes.
Maybe just, like, quick story.
Can you just talk about that?
That was just really fun.
Ether Zero was...
a hell of a project because conceptually it was a very short project of like, hey, people have made a lot of progress in verifiable rewards in math and in code.
Let's say we do it in chemistry.
So chemistry is like not a verifiable field, right?
Like, of course, you can go test something in the lab, but then we like had to think about all these like ways that we make chemistry verifiable.
And one of the ones we settled on was like, make a molecule that has like three nitrogens, two oxygens, 10 hydrogens or something.
And we thought that was like a pretty verifiable.
Pretty verifiable question.
But every time we would train a model, it would find some new, insanely weird trick to generate these molecules.
And I'll just tell you, one of the examples was that it would make these molecules and we would do some checks to make sure it had the right bonds, the right number of electrons, the right number of atoms and stuff like that.
But it would just solve the problem in any way possible.
So it would just put all the nitrogens over here, put all the oxygens over here, just things that don't look good.
So we started coming up with these rules of like, oh, let's check.
to make sure it followed these good practices or these good practices.
And we found ourselves into this, like, you know, it's like the opposite of the bitter lesson, like, I don't know, the boutique lesson where you like try to make everything custom.
But one of the things it kept doing is it kept putting these nitrogens in a row and it put like one nitrogen, two nitrogen, three nitrogen, all in a chain.
And this is like, you know, if you have three nitrogens, it's like explosive, you know, two nitrogens is like bad and like four nitrogens you can't make.
And I kept telling everyone, like it would make these like six nitrogen compounds and they're just literally impossible and they're not possible.
Many of the people on the team were like computer scientists like on this team.
And one of them like one day sent me that like this is on the cover of Nature Today on Nature's website.
Somebody made a six nitrogen compound.
And this is like somebody's like career.
to deliver this compound because this is the most unstable, like insane compound.
You can make some ridiculous setup and like the spectroscopy to get that proven was like very difficult.
I don't know how they did it.
It was amazing accomplishment.
Like, look, Andrew, like it's not actually impossible.
And it was so funny to me that like our model was sitting here spitting out these six nitrogen compounds in like, you know, 2024.
or 2025.
And like the paper just happened to come out that year that like mankind had finally made a six nitrogen compound.
Do you think that those were actually synthesizable even under these extreme circumstances?
No, no.
Our model was just, it was just reward hacking.
Okay.
And it was just, the model was so creative and ways to reward hack.
Like one of the, another one we did was, you know, we wanted it to make sure that when it proposed a reaction, like make this compound, tell me how to make this compound.
We would try to make it.
sure that all the reagents were purchasable like you could purchase them they were not like made up yeah um and the reason we came up with that is that originally we just like take the end compound and then like remove one atom and be like here's you buy this and then put the atom on it's like okay it's like well i wish it was like that um so they have to be purchasable and then we'll be like we thought it might be hard if they're all purchasable because sometimes you actually order things custom or something so we'll just make sure one purchasable So the first thing it starts doing is putting nitrogen in there because nitrogen is purchasable and it like has no participation in the reaction, right?
Like, oh my God.
Okay.
So then like, okay, it has to be purchasable.
It has to participate in the reaction.
Then it starts putting like acid base chemistry.
We'll just put an acid here.
Acids are purchasable and it'll move one atom.
And they're like, okay, fine.
Can't be that.
Everything has to be purchasable.
Then we find ourselves and I'm like sitting there one day building this like ridiculous.
catalog of purchasable compounds and a bloom filter so it can go fast enough on our training loop and i'm like why am i doing this how did i get here and i don't know it was really funny because um pre-training or training transformers you know on just data like just supervised training where you just have the inputs and the outputs directly very nice relaxing you know like things are always robust you know things go pretty smoothly when we do these verifiable rewards where you have to like write a bulletproof verifier it is really difficult.
And we had so many models trained only to find out they were hacking some other like random thing in our setup.
It's really hard.
And I, and I, I don't envy the frontier labs that have to do this at a very massive scale because we had a lot of adventures in ether zero and you guys should read the blog post.
It's very fun.
GRPO.
We did make some modifications to GRPO.
Yeah.
I actually, I used to know.
all the names of these modifications but uh i think it's like uh dappo is one modification and like the clipping we did was special and we explored a lot of that stuff yeah um and it was uh also one of these things where like you think the hypers are wrong the algorithm is wrong and then you find out it's just because like you had You somehow sorted the reagents when you made your training data, but when you made your test data, you didn't sort them alphabetically.
And the model was just like barfing because its whole strategy was to exploit something in the way you sort of things.
So, yeah, we explored a lot of different methods.
And I learned a lot about chemistry, a lot about nomenclature.
And actually, I learned a lot about medicinal chemistry as well, more than I ever wanted to.
Awesome.
If you want to do some like engineering, just check out Edison Scientific.
And they have, you know, I think a lot, they're hiring with lots of like interesting things, everything from scientists to, you know, infrastructure engineer.
Yeah.
Thanks, Andrew, again.
Yeah.
Thank you very much for joining us.
