Gaussian splatting killed my mesh nostalgia
3D Gaussian Splatting represents a scene as millions of fuzzy, colored blobs and renders it in real time at NeRF quality. It does this without a neural network.
The paper everyone in 3D graphics is talking about this month is “3D Gaussian Splatting for Real-Time Radiance Field Rendering,” by Kerbl, Kopanas, Leimkühler and Drettakis, presented at SIGGRAPH a couple of weeks ago. It matches or beats the quality of the best NeRF methods for captured scenes, trains in tens of minutes, and renders at over 100 frames per second at 1080p. It uses no neural network at all.
I’ve followed this line since the original NeRF in 2020 and Instant NGP last year. Each step moved more of the representation out of the network and into an explicit structure. Gaussian splatting finishes that move.
A scene is represented as a few million 3D Gaussians. Each one is a small, fuzzy ellipsoid with a position, a 3D covariance (its shape and orientation, stored as a scale and a rotation), an opacity, and a color that can change with viewing direction, stored as spherical harmonics coefficients so shiny surfaces look right from different angles.
To render, you project each Gaussian onto the screen, where it becomes a 2D ellipse, sort them by depth, and blend them front to back, which is basically alpha compositing. They wrote a fast tile-based rasterizer on the GPU that does this for millions of splats in real time. It’s differentiable, so training is the same idea as NeRF: render the training views, compare with the photos, and update every Gaussian’s parameters by gradient descent. The training also adds Gaussians where the scene is under-reconstructed and removes transparent ones, so the representation grows detail where it’s needed.
It starts from a sparse point cloud produced by structure from motion, the same camera calibration step NeRF needs, so the Gaussians start roughly in the right places.
I like being able to inspect the scene directly. Every piece of the scene is a primitive you can see, count, move and delete. That’s what I loved about meshes in grad school and what NeRF gave up. And it’s fast on normal graphics hardware, because it’s rasterization, which GPUs have been built for since the 90s. But it also keeps what NeRF got right: soft, view-dependent appearance, learned from photos, that captures detail no artist would model by hand.
For a decade I’ve half-expected learned representations to replace explicit geometry entirely. This paper suggests a different outcome: explicit primitives that are optimized like a neural network. It’s not a mesh, and it’s not a network either.
My guesses: phone apps that capture a splat scene in minutes within a year. Game engines and web viewers adding support. And then the hard part, which is the same one as always for me: making these scenes move. Animated splats, splat avatars, faces made of Gaussians. And generated splats, whole 3D worlds from a prompt or a single image, which would turn generative models from making pictures into making places.