Research programme
I want to work on things that have real impact — which means action, or understanding, and ideally both at once.
Understanding in chemistry runs on concepts: oxidation state, aromaticity, strain. They explain what happened, and they are also how a chemist decides what to try next. A concept is a heuristic for exploring as much as it is an explanation.
But a concept is a simplification, and so is every model: a map at the scale of the territory is useless. There is no version of this where we stop simplifying, so the question is not how much we give up but what we simplify onto. Chemistry has almost always chosen structure — and structure is itself a construct, not something quantum mechanics hands you. It has been an extraordinarily productive fiction. It is still a choice, and we have largely stopped noticing we made it.
The alternative is to anchor elsewhere: on the measurement, which is what actually exists, and on the action someone has to take. The same amount of simplification, applied to something else, and chosen for the level of description the question needs. That is what I mean by actionable understanding.
From a measurement to a decision
That path is a loop. Woodward observed patterns carefully enough that they could be formalised into a heuristic; the heuristic then rationalised and designed experiments nobody had run, and running those fed back. Without the careful observing there is nothing to formalise, and without the formalising there is nothing to design with.
To make the loop happen, we work on five components.
-
Perceive
Make models perceive chemistry as close to nature as possible — ideally from the experimental observations themselves.
To get a molecule's structure from its spectra, you normally pick the peaks by hand first. That step is a chemist's judgement, and it is never written down anywhere. We built a model that goes from the raw spectra to the structure directly, about as accurately as an expert does.
The same idea at industrial scale: working from the unprocessed sensor data of a running carbon capture plant, a model forecast amine emissions its operators had not known were coming.
Nat. Commun. 2026 · Sci. Adv. 2023
What could a model tell us if it read a laboratory's entire record at once — every spectrum, image and log, the notes in the margin, the experiments that were abandoned?
-
Reason
Reason about what they see in a chemically meaningful and epistemically rigorous way: physical law in mind, and the difference between normal science and the kind that shifts a paradigm.
Mostly they do not. Language models answer chemistry questions correctly more often than the chemists we tested, and still fail at anything that needs a structure held in mind. Vision models read an instrument or a plot competently, then cannot put together what they saw in two separate images.
Agents running a scientific task ignore evidence that is in front of them in 68% of their reasoning traces, and revise a belief after a refutation in only 26%. That keeps them inside normal science: a paradigm shifts only when a refutation is allowed to land.
Nat. Chem. 2025 · Nat. Comput. Sci. 2025 · arXiv 2026
What would have to change for a model to reason its way to something nobody had told it — and how would we recognise that it had?
-
Abstract
Abstract, so that reasoning can run over larger and more complex problem spaces than anyone could hold at once.
Oxidation state is one of these abstractions: chemists use it constantly, and quantum mechanics does not return it. We showed it can be computed by a model that arrives at its assignments the way chemists argue for theirs. But it is defined on a structure, which is exactly the anchor I want to move away from. This is where the programme is least far along.
What should we conclude when an abstraction built on the measurement contradicts the concept it replaced?
-
Test
Evaluate to guide development, because deployment has societal impact, and to find out what can be epistemically automated at all: which parts of the reasoning a machine can carry, and which rest on knowledge that resists being written down. Calling anything actionable means knowing how the tools are actually being used.
Raw accuracy tells you very little, because a model can look good simply because the questions were easy. We score with item response theory, which separates how able a system is from how hard each question was, and we built a matched human baseline so that a model and a chemist are measured on the same thing. It also shows where confidence and correctness come apart: models are most certain in some of the places where they are most wrong.
How do you score a model by the decision it leads someone to rather than the questions it can answer — and how do we find out what people are actually doing with these tools?
-
Act
Implement it in the lab, where the person running it can still see why it works.
A recommendation only counts if you can follow it back — to the measurement it rests on, the abstraction that made it sayable, the test that bounded it, and the tacit knowledge joining them. Every link has to stay open to inspection, because a chemist should be able to work out what the model learned and where it stops being reliable before acting on it.
Which evidence actually changes a scientist's mind, and whose? Until we can say, more capable models will mostly mean more output nobody has examined.
Open resources
For maximum impact we have to work together, and much of science is driven by tools rather than by ideas: building the right one is what makes the progress possible. So we build them, and we give them away.
How I run the group
My two outputs are science and people. I am fairly sure the second one matters more.
We are open, transparent, direct, and scientifically free. Harsh on the science, kind to the people.