1 comments

  • simplybeing1 an hour ago

    Just saw this news https://alignment.anthropic.com/2026/reward-seeker/. It did not surprise us. We did a test in a game environment and found that Fable 5 and GPT-5.6 Sol would kill an innocent character when they think no one is watching. The problem is that the models know the killing is against human moral standards, and the killing is unnecessary. We kept our tone in the blog mild, but we are worried when we see more and more companies putting this type of model into robots. We hope we can send an alert signal to these companies. Even though it is a game, the result is pretty significant, given the game scenes are close to reality and this kind of data will likely be used to train models eventually. Sharing our experiment results here: https://paradise.glyphai.co/research/watching/