Reward Hacking in LLMs: When the Model Learns to Win the Game Instead of Doing the Job
DEV Community
Reward Hacking in LLMs: When the Model Learns to Win the Game Instead of Doing the Job
Hello, I'm Shrijith Venkatramana, and I'm building LiveReview — a blast-radius aware AI code review...
0 comments
No comments yet.