AlphaFold and learning physics from data
DeepMind's AlphaFold 2 predicts protein structures about as accurately as experiments do. A fifty-year-old physics problem, solved mostly by learning from examples.
On Monday the organizers of CASP, the competition that has benchmarked protein structure prediction every two years since 1994, announced that DeepMind’s AlphaFold 2 had essentially solved the problem as the competition defines it. Its median accuracy score across the hardest targets was around 92 on a scale of 100. Scores around 90 are considered competitive with experimental methods like X-ray crystallography, which can take months or years per protein.
I’m not a biologist, but the protein folding problem is one of those things every engineering student hears about. A protein is a chain of amino acids. The sequence is easy to read from DNA. The chain then folds into a specific 3D shape, and the shape determines what the protein does. Predicting the shape from the sequence has been an open problem for about fifty years. In principle it’s pure physics: atoms, bonds, electrostatic forces, water. In practice, simulating that physics directly is far too expensive for all but the smallest proteins, and the search space of possible folds is astronomically large.
DeepMind hasn’t published the full method yet, only a blog post and the CASP abstract. What they’ve said is that it’s an attention-based network trained end to end on the roughly 170,000 known structures in the Protein Data Bank, using the sequences of related proteins from evolution as a major input. Amino acids that mutate together across species tend to be in contact in the folded structure, and the network learns to read that signal. It iterates, refining a guess of the structure, and outputs 3D coordinates with its own confidence estimates.
As someone with a mechanical engineering background, what strikes me is how little of the physics is written into it, at least as far as we know now. The traditional approach is to model the forces and search for the lowest-energy fold. AlphaFold instead learns from examples of what folded proteins look like, and from evolution’s record of which parts belong together. The physics is in the data. The model found a shortcut to the answer that physics would eventually give, without simulating the physics.
I’ve been saying that the next decade is about machines understanding the physical world. I was thinking about robots and bodies. AlphaFold is a different angle on the same claim. A learned model can capture the behavior of a physical system well enough to replace simulation, if you have enough good examples of the system’s outcomes. That’s both hopeful and a little humbling for anyone who spent years writing equations of motion.
The question I’d love answered is whether the model has learned something like physics or something like a very good lookup of evolutionary patterns. It probably doesn’t matter for biologists, who just got a tool that will change their field. It matters a lot for whether this approach works in places where there’s no evolution to learn from.