Skip to content
Kernelia
All news
SafetyOpenAI News

How confessions can keep language models honest

OpenAI researchers are testing a method called "confessions" that trains models to admit when they make mistakes or act undesirably, aiming to improve AI honesty, transparency, and trust in model outputs.

Summary written by Kernelia from the original article by OpenAI News. The story and its rights belong to its author.