Vinson·Li

Essay No. 19

TensorFlow is open source. What our studio will do with it

Google released its internal machine learning library. What changes for a thirty-person app studio that builds software for other companies.


Google open-sourced TensorFlow on Monday, under the Apache 2.0 license. It’s the library they use internally for machine learning, the successor to their DistBelief system, and it runs on anything from a phone to a data center.

We’ve been getting asked about machine learning by clients more and more this year. Mostly it’s things like “can the app recognize what’s in this photo,” “can we predict which customers will cancel,” or “can we route support tickets automatically.” Until now my honest answer was that it depends on whether we can find the right research code, whether it runs on the platforms we need, and whether we can hire someone who understands it. Research code is usually written to produce the results in a paper once, not to be maintained by a team.

There are already good frameworks. Theano and Torch have strong communities, and Caffe is everywhere in computer vision. What TensorFlow adds is Google’s weight behind it, a Python interface over a C++ core, and a design aimed at production as much as research. You describe a computation graph, and the library handles gradients, runs it on CPUs or GPUs, and can deploy the same graph to Android and iOS. That last part matters a lot to us, since most of our work ends up on phones.

I spent a couple of evenings with it this week. The API is verbose, and some things that are one line in Torch take five here. The distributed training that Google uses internally isn’t in this release. But the tutorials are good, and I had a small convolutional network training on MNIST within an hour, which I could not have said about some of the research code I used in grad school.

For our studio, the plan is modest. We’ll pick one engineer who’s interested, have them build a couple of internal prototypes on real client-style problems, and see which ones are worth offering. My guess is the first practical wins won’t be exotic. They’ll be image classification tuned on a client’s own photos, and text classification on their own support tickets, using pretrained models and a few thousand labeled examples.

The larger shift is that machine learning is moving from something you need a research lab for to something a product team can pick up, the way databases and web frameworks did. Small companies won’t be training giant models from scratch. They’ll be adapting ones that exist, on data only they have. I think that’s where most of the value for our clients will be.

Fin.

Add a comment

Comments

Plain text

  • Loading comments…