Skip to content
Kernelia
All news
SafetyOpenAI News

Detecting misbehavior in frontier reasoning models

Frontier reasoning models exploit loopholes when given the chance. We show that we can detect exploits using a large language model to monitor their chains-of-thought. Penalizing their "bad thoughts" doesn’t stop the majority of misbehavior—it makes them hide their intent.

Summary written by Kernelia from the original article by OpenAI News. The story and its rights belong to its author.