Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Note that this reinforcement finetuning is something different than regular RLHF/DPO post training


Is it? We have no idea.


Yes it is. In RLHF and DPO you are optimizing the model output for human preferences. In the reinforcement fine tuning that was announced today you are optimizing the hidden chain of thought to arrive at a correct answer, as judged by a predefined grader.


I mean i think it could easily be PPO post training. if your point is that the rewards are different, sure




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: