SafetyOpenAI News
How confessions can keep language models honest
OpenAI researchers are testing a method called "confessions" that trains models to admit when they make mistakes or act undesirably, aiming to improve AI honesty, transparency, and trust in model outputs.
Summary written by Kernelia from the original article by OpenAI News. The story and its rights belong to its author.

