If it truly is the training data that's making models smart, then that would explain that there is both a minimum and maximum "useful" size to LLMs. The recent stream of papers seems to indicate that the cleaner the input data, the less size is required.
That would negate, at least partially, the "we have 20 datacenters" advantage.
I think it is. I also think this is what OpenAI did. They’ve carefully crafted the data.
I don’t think they have an ensemble of 8 models. First, this is not elegant. Second, I don’t see how this could be compatible with the streaming output.
I’d guess that GPT4 is around 200B parameters, and it’s trained on a dataset made with love, that goes from Lorem Ipsum to a doctorate degree. Love is all you need ;)
It's still large. But it might no longer need to be "only few entities on the planet can afford to make one, and not at the same time, since NVIDIA can pump out GPUs only so fast" large.
That would negate, at least partially, the "we have 20 datacenters" advantage.