There was steady improvement for decades (e.g. on a Mac, you can compare Fred, Victoria, Vicki, and Alex, as representatives of 4 generations of TTS, and there were year-over-year improvements as well).
But the latest neural techniques have added quite a bit of naturalness, and their computational requirements, while high, are within the reach of consumer level devices.
But the latest neural techniques have added quite a bit of naturalness, and their computational requirements, while high, are within the reach of consumer level devices.