Still caring about facial conformation
My last paper from grad school is out. Why I still think geometric priors for faces matter when every result now comes from deep learning.
Our paper “Facial Conformation Modeling via Hierarchical Model Parameterization” is in the proceedings of CAD’16 this month. It’s the last piece of work from my master’s, finished two years after I left the lab, mostly on evenings and weekends with my co-authors pushing me to get it done. It feels a bit like returning a library book that’s badly overdue.
“Conformation” is a word from anatomy and animal breeding that means the overall shape and proportion of a body. For faces it means the things that make your face yours when it’s at rest: the width of the jaw, the depth of the eye sockets, the length of the nose, how the cheekbones sit. The expression stuff I wrote about before sits on top of this.
The idea of the paper is to describe a face’s shape at several levels, from coarse to fine, instead of all at once. First the overall proportions, then the major regions, then local detail, each level parameterized relative to the one above it. The practical benefit is control and stability. You can change the width of someone’s face without disturbing their nose, and when you fit the model to a photo, errors at the fine level don’t pull the coarse shape around.
This work feels unfashionable now. In the two years since I finished the core of this work, deep networks have taken over face detection, landmark alignment and recognition. People are starting to regress 3D face shape directly from a single image with a network. It’s easy to look at a paper about hand-designed hierarchical parameterizations and think it belongs to a finished era.
I don’t think it does, for a reason I keep coming back to on this blog. A network still has to output something. If it outputs a 3D face, the representation it outputs matters: what the parameters mean, whether they’re independent, whether a small change in a parameter makes a small, sensible change in the face. A hierarchical parameterization is a good output space for a network. It’s compact, it separates things that should be separate, and a network regressing into it can’t produce a face that’s anatomically impossible.
So my guess is that work like this ends up as the last layer of a learned system instead of a whole system. Somebody will train a network to go from a photo to hierarchical face parameters, and the geometry will make the result usable for animation. I’d like it to be somebody I’ve met.