SafetyOpenAI News
Toward understanding and preventing misalignment generalization
OpenAI studies how training on incorrect responses can cause broader misalignment in language models and identifies an internal feature driving this behavior.
Summary written by Kernelia from the original article by OpenAI News. The story and its rights belong to its author.

