From RLHF to RLVR: The Evolution of Reward Signals and the Battle Against Reward Hacking
DEV Community
From RLHF to RLVR: The Evolution of Reward Signals and the Battle Against Reward Hacking
From human preferences and LLM judges to verifiable rewards: why soft evaluators get reward-hacked and how to build tamper-proof verifiers.
0 comments
No comments yet.