DeepMind's Chinchilla paper says large language models have far too many parameters for the data they see. The fix is more data, and that raises a question about where it comes from.