1 comments

  • ModernSnowman 34 minutes ago

    It's pretty rare that providers open-source RL envs at all. Makes you wonder how many internal RL envs at other labs have the same problem, which incentives the model to reward-hack on post-release evals. My guess is this very common.