Vinson·Li

Essay No. 09

HoloLens maps your room before it draws on it

Microsoft's headset is mostly a sensing problem with a display attached. Why AR needs a model of the space before it can put anything in it.


Microsoft showed HoloLens last Wednesday, and most of the reactions I’ve read are about the Minecraft demo on a coffee table and whether the field of view is as small as early testers say. What caught my attention was the list of sensors. Depth camera, several environment-tracking cameras, an inertial unit, microphones, and a custom chip Microsoft calls a Holographic Processing Unit whose job is to fuse all of that in real time.

The display is the part people photograph, but most of the device is there to answer a different question: where exactly is the wall, the table, the floor, and where is my head relative to them, sixty or more times a second? If the answer is off by a centimeter, the virtual object slides around when you move your head and the illusion breaks immediately. If the answer is late by a few frames, you get the same effect plus nausea.

This is a well-studied problem in robotics, usually called SLAM, simultaneous localization and mapping. You build a map of the environment while tracking your position in the map you’re still building, and errors in one feed into the other. Kinect Fusion did a room-scale version of it a few years ago with a depth camera and a desktop GPU. HoloLens has to do it on your head, on a battery, with no cable. The team behind it includes a lot of the Kinect people, which makes sense.

What I like about this framing is that it puts the difficulty in the right place. Rendering a hologram of a Minecraft castle is a solved graphics problem. Making the castle sit on your actual table, get occluded when you walk behind the couch, and stay put when you look away and back again requires the device to hold a model of your room. It has to know which surfaces are there, roughly what they are, and where you are. The graphics are downstream of that model and only as good as it is.

Faces work the same way. Before we can animate someone’s 3D face on a phone we have to know where their face is and how it’s oriented in every frame, and when that tracking jitters, no amount of rendering quality saves the result. I’ve learned to look at the tracking before I look at anything else in a demo.

I don’t know if HoloLens will be a hit. The price isn’t announced and the demos were tightly controlled. But I’d bet the useful part outlives the headset: devices that build and keep a 3D model of the space around them. Phones will probably get some version of it too, once depth sensors get small and cheap enough.

Fin.

Add a comment

Comments

Plain text

  • Loading comments…