Graduated. My thesis in one paragraph, and how deep learning will eat it
What my master's thesis on 3D face meshes does, which part of it I think neural networks will replace soon, and which part I think they won't.
Finished my M.Eng at UBC this month. The thesis title is “triangle mesh parameterization for 3D face modeling and animation,” and my family still doesn’t know what it means, so I’ll try.
You have a template face, a mesh of a few thousand triangles that someone rigged once for animation (controls for the mouth, brows, eyelids and so on). When a new person comes along, you deform the template until it fits their face, and the correspondence has to be right: the template’s left mouth corner lands on their left mouth corner, the eyelid on the eyelid, and everything between follows smoothly. Get that right and the rig works on the new face, so you can animate a stranger without rigging them from scratch. Most of my two years went into making that mapping stable enough that the animation doesn’t break.
I did mechanical engineering in Wuhan before this, and more of it carried over than I expected. A triangle mesh is close to a finite element mesh, and bending one so that certain points hit exact targets is an old problem for anyone who has worked with plates and shells. Tolerance is where they differ. Nobody can see that a bracket is off by 2%, but when a face is off by that much people notice right away.
The weakest part of my pipeline is the front, where I need a few dozen landmarks on the photo (mouth corners, eye corners, nose tip, jaw line). Right now that’s a hand-built detector with engineered features and a statistical shape model, iterated until the points fit. It’s fine in good light and drifts in a dim restaurant, which is unfortunately where people take selfies.
I think learned models replace this part soon. AlexNet cut ImageNet error roughly in half in 2012, Google bought DeepMind in January for reportedly several hundred million dollars, and in March Facebook’s DeepFace paper got 97.35% on Labeled Faces in the Wild, about where humans are. I’d guess that in two or three years nobody hand-tunes landmark detectors anymore, and not long after that networks go from one photo to a 3D face with no explicit landmarks at all. Some of what I derived by hand over two years will become a loss function and a weekend of GPU time. I have mixed feelings about that, but mostly I’m glad I got to see it coming from inside the problem.
What I’m less sure about is the output. A network that predicts a 3D face still has to output a depth map, a point cloud, coefficients of some template, something. If it outputs a mesh with consistent correspondence (vertex 1,204 is the left mouth corner on everyone) you can animate it and transfer expressions between people. A depth map with the same benchmark score can’t be animated at all. So I think the geometry half of my thesis stays useful for a while even after the detection half is gone.
For now I’m still working on the app I started last April, which turns one selfie into an animated 3D face on the phone, and doing some work for a small app studio downtown. On the app, the lesson so far is that failure cases matter more than average accuracy. When the landmarks are wrong the face looks horrifying, and nobody writes to tell us, they just uninstall.