A whole scene stored inside a network
NeRF represents a 3D scene as a small neural network you can render from any viewpoint. No mesh, no triangles. My old thesis problem, answered from a different direction.
I finally read the NeRF paper properly this week, “Representing Scenes as Neural Radiance Fields for View Synthesis,” by Mildenhall and others from Berkeley and Google. It came out in March and got an oral at ECCV in August, and follow-up papers are appearing every week. I’ve been putting it off because I suspected it would make me rethink things I thought I understood about 3D. It did.
The input is a set of photos of a scene, maybe a hundred, taken from known camera positions. The output is a small neural network, a plain MLP with no convolutions, that takes a 3D position and a viewing direction and returns a color and a density. Density means how much stuff is at that point. Color can depend on the viewing direction, which lets it capture reflections and shiny surfaces.
To render an image from a new viewpoint, you shoot a ray through each pixel, sample points along the ray, ask the network for color and density at each point, and composite them front to back with the standard volume rendering equation. Because every step is differentiable, you can train the network by rendering the training viewpoints and comparing to the actual photos. After training, the scene lives in the network’s weights. The results are startling: fine detail, thin structures, reflections, all rendered from viewpoints nobody photographed.
Two tricks make it work. One is positional encoding: before feeding coordinates into the network, they map each one through sines and cosines at many frequencies. Without that, MLPs produce blurry results because they’re biased toward smooth functions. It’s the same idea the Transformer paper used for word positions. The other is hierarchical sampling, which spends more samples along each ray where the density actually is.
I spent my master’s making triangle meshes behave. A mesh is an explicit representation: a list of vertices and faces you can edit, animate and render fast on any GPU. NeRF is implicit. There’s no surface anywhere, just a function you can query. That gives up almost everything meshes are good at. It’s slow to render, takes hours to train per scene, can’t easily be edited or animated, and it’s one network per static scene. In exchange, it captures appearance with a realism I never got close to with meshes.
My guess is that neither side simply wins. Someone will find a representation that keeps NeRF’s quality and trains in minutes instead of hours, probably by putting more of the information into a data structure and less into the network. And someone will find a way to make these fields move, which is where faces come back in. A NeRF of a face that you could animate with expressions would be the thing I was trying to build in 2013, done properly.