Skip to content
Kernelia
All news
SafetyOpenAI News

The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions

Today's LLMs are susceptible to prompt injections, jailbreaks, and other attacks that allow adversaries to overwrite a model's original instructions with their own malicious prompts.

Summary written by Kernelia from the original article by OpenAI News. The story and its rights belong to its author.