Skip to content
Kernelia
All news
SafetyOpenAI News

Toward understanding and preventing misalignment generalization

OpenAI studies how training on incorrect responses can cause broader misalignment in language models and identifies an internal feature driving this behavior.

Summary written by Kernelia from the original article by OpenAI News. The story and its rights belong to its author.