NeRF went from hours to seconds
Nvidia's Instant NGP trains a neural radiance field in seconds with a multiresolution hash table. The representation mattered more than the network.
In October 2020 I wrote that someone would find a way to keep NeRF’s quality and train in minutes instead of hours, “probably by putting more of the information into a data structure and less into the network.” Last week Nvidia posted “Instant Neural Graphics Primitives with a Multiresolution Hash Encoding,” and it trains a NeRF in about five seconds to a minute on one GPU. I hadn’t expected it to get down to seconds so soon.
The original NeRF stores the whole scene in the weights of an MLP with eight layers of 256 units. Every query, a position and a direction, runs through the full network. Training has to adjust all those weights to fit every detail of the scene, which takes a long time, and every rendered pixel needs hundreds of network queries.
Instant NGP moves most of the scene into a lookup table. Picture the scene covered by grids at many resolutions, from coarse to very fine. At each grid level, the corners of the cell containing your query point each have a small learned feature vector. You interpolate between the corners at each level, concatenate the results from all levels, and feed that into a tiny MLP, just a couple of layers, which outputs color and density.
The trick is the hash. A fine 3D grid would have far too many corners to store. So each level maps its corners into a fixed-size table using a spatial hash, and many corners share the same slot. Collisions happen, and the network is left to sort them out. Since most of space is empty, and the gradients from important regions dominate, the table ends up storing what matters. The memory is fixed and the lookup is cheap. They also wrote everything as fused CUDA kernels, which is a large part of the speed.
What I take from it is the same lesson as my thesis, in a new form. The quality and cost of a system depend enormously on the representation you choose. NeRF showed that a neural representation could capture scenes beautifully. Instant NGP shows that most of the work of that representation belongs in a smart data structure, with a very small network to decode it. The learning is still there. It’s just been moved to where it’s cheap.
A NeRF in seconds changes what’s possible. Capturing a room with a phone and walking around it in 3D a minute later is no longer research. I’d expect 3D capture apps built on this within a year, and more work on making these fields editable and dynamic, which is the part that interests me most. A static scene is a photograph you can walk into. A scene you can change is a world.