Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Over yonder: https://x.com/alexalbert__/status/1780707227130863674

my $0.02: it makes me very uncomfortable that people misunderstand LLMs enough to even think this is possible



It is 100% possible for performance regressions to occur by changing the model pipeline and not the model itself. A system prompt is a part of said pipeline.

Prompt engineering is surprisingly fragile.


Absolutely! That was covered in the tweet link. If you're suggesting they're lying*, I'm happy to extract it and check.

* I don't think you are! I've looked up to you a lot over last year on LLMs btw, just vagaries of online communication, can't tell if you're ignoring the tweet & introducing me to idea of system prompts, or you're suspicious it changed recently. (in which case, I would want to show off my ability to extract system prompt to senpai :)


I was agreeing with the tweet and think Anthropic is being honest, my comment was more for posterity since not many people know the difference between models and pipelines.

Thanks for liking my work! :)


Is that surprising? Seemed like a giant hack to me. Prompt engineering sure sounds better than hack though.


It is a necessary hack, though.


Of course it is possible. For example via quantization. Unless you are refering to something I can't see in that tweet. (not signed in).


You're right, that's a good point. It is possible to make a model dumber via quantization.

But even F16 -> llama.cpp Q4 (3.8 bits) has negligible perplexity loss.

Theoratically, a leading AI lab could quantize absurdly poorly after the initial release where they know they're going to have huge usage.

Theoratically, they could be lying even though they said nothing changed.

At that point, I don't think there's anything to talk about. I agree both of those things are theoratically possible. But it would be very unusual, 2 colossal screwups, then active lying, with many observers not leaking a word.


Thanks, this is the tweet thread I was referring to.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: