Vinson·Li

Essay No. 64

Why our models train on forty different lobbies

A year of live events taught us that benchmark accuracy barely predicts real-world accuracy. What we changed in how we collect data and evaluate.


Last September I wrote about our first live conference check-ins and how most of our problems came from lighting, queues and printers. A year later, after many more events, I’d sum up what we learned in one sentence: benchmark accuracy barely predicts how the system performs in a real lobby.

When we started, we evaluated our models the way the field does. We used public face benchmarks and our own test sets of well-lit, front-facing photos. Our numbers were excellent, as are everyone’s. In the field, what mattered was different: people at an angle, backlit by a window, under colored stage lighting, wearing a hat, holding a coffee, or with a new haircut since their registration photo two months ago. None of that shows up much in benchmarks.

So we changed how we work.

We evaluate per venue, not globally. After every event, with consent and within our retention policy, we look at aggregate metrics for that venue: match rate, how often people needed a second attempt, how often they gave up and went to the desk. We’ve now done this across dozens of venues. A model change that improves the global average but hurts the three worst venues doesn’t ship.

We build test sets around conditions, not people. Instead of “a thousand faces,” our internal evaluation sets are organized by situation: strong backlight, low light, side angle, glasses glare, hats, masks for costume events. Each has its own threshold. That makes it much easier to see where a new model got better or worse, and it stops the average from hiding the tails.

We treat the environment as part of the product. Kiosk placement, the angle of the iPad, a small light we ship with every kit, and the on-screen prompt all changed results more than most model updates. We have a setup checklist per venue now. It isn’t exciting work, and it’s probably our biggest improvement this year.

The general lesson I’d give anyone deploying models in the physical world: your real test set is the world, and it’s shaped very differently from your benchmark. Collect variety on purpose, measure the worst cases separately, and don’t let a good average convince you you’re done. Our worst lobby taught us more than our best benchmark.

Fin.

Add a comment

Comments

Plain text

  • Loading comments…