Datasets have a consent debt
Microsoft quietly took down MS-Celeb-1M, ten million photos of a hundred thousand people. A lot of the face recognition industry was built on data nobody agreed to give.
The Financial Times reported last week that Microsoft has deleted MS-Celeb-1M, a dataset of about 10 million face images of roughly 100,000 people, gathered from the web. Duke and Stanford took down two other face datasets, one from campus surveillance cameras and one from a café webcam, around the same time. Microsoft’s statement said the site was intended for academic purposes and was run by an employee who is no longer at the company.
MS-Celeb was nominally a list of “celebrities,” but the researcher Adam Harvey, who has been cataloguing these datasets, found that it included plenty of people who aren’t famous in any normal sense: journalists, activists, academics, writers who happen to have a web presence. None of them were asked. The data was used widely, by university researchers and by companies, including some Chinese companies building surveillance products.
I’ve been in this field since before deep learning took it over, and I’ll say the uncomfortable part: this is how the whole industry was built. Face recognition got good between 2014 and 2017 because networks got better and because people assembled face datasets with millions of images scraped from the web. Almost none of those people agreed to be training data. It was just how things were done. Everyone downloaded the same datasets and compared numbers on the same benchmarks.
I think of it as a consent debt. The industry borrowed against the privacy of millions of people to build its models, and that debt is now coming due, in the form of deleted datasets, lawsuits under laws like Illinois’s biometric privacy act, and bans like San Francisco’s last month. Models trained on those datasets are still out there, running in products, and deleting the dataset doesn’t un-train them.
At Amanda we’ve tried to be careful about what we use and why, but it would be dishonest to pretend any company in this field is entirely separate from the history. The pretrained models and benchmarks everyone relies on came from somewhere.
What would paying down the debt look like? A few things seem clear to me. New datasets should be collected with consent and documentation, even though that’s slower and more expensive. Synthetic faces, which StyleGAN made much more realistic, could cover a lot of what scraped data was used for, at least for testing. And companies should be able to say where their training data came from. If you can’t answer that question for your model, you’ll probably have to answer it later in a less pleasant setting.