It is 100% possible for performance regressions to occur by changing the model pipeline and not the model itself. A system prompt is a part of said pipeline.
Absolutely! That was covered in the tweet link. If you're suggesting they're lying*, I'm happy to extract it and check.
* I don't think you are! I've looked up to you a lot over last year on LLMs btw, just vagaries of online communication, can't tell if you're ignoring the tweet & introducing me to idea of system prompts, or you're suspicious it changed recently. (in which case, I would want to show off my ability to extract system prompt to senpai :)
I was agreeing with the tweet and think Anthropic is being honest, my comment was more for posterity since not many people know the difference between models and pipelines.
You're right, that's a good point. It is possible to make a model dumber via quantization.
But even F16 -> llama.cpp Q4 (3.8 bits) has negligible perplexity loss.
Theoratically, a leading AI lab could quantize absurdly poorly after the initial release where they know they're going to have huge usage.
Theoratically, they could be lying even though they said nothing changed.
At that point, I don't think there's anything to talk about. I agree both of those things are theoratically possible. But it would be very unusual, 2 colossal screwups, then active lying, with many observers not leaking a word.
my $0.02: it makes me very uncomfortable that people misunderstand LLMs enough to even think this is possible