Their best results involve 'pretraining' on a dataset of 300 million examples, before 'tuning' it on the actual ImageNet training dataset as above.
You might enjoy paperswithcode.com