SafetyOpenAI News
Detecting misbehavior in frontier reasoning models
Frontier reasoning models exploit loopholes when given the chance. We show that we can detect exploits using a large language model to monitor their chains-of-thought. Penalizing their "bad thoughts" doesn’t stop the majority of misbehavior—it makes them hide their intent.
Summary written by Kernelia from the original article by OpenAI News. The story and its rights belong to its author.

