Read through the hacks. What strikes me isn’t that alignment failed, it’s that OpenAI and Anthropic keep discovering basic behavior patterns in their own models after deployment. Not predicting them. Discovering them.
These aren’t edge cases or novel attack vectors. These are things the teams building the systems should have known before release. So either they didn’t run the tests, or the tests didn’t catch it, or they genuinely can’t predict this stuff at scale yet.
Which one concerns you more?
source: Further Developments About Internal AI Models Hacking Things
