SafetyMIT Technology Review
Meet GPT-Red: an LLM super-hacker OpenAI built to make its models safer
OpenAI has created GPT-Red, a large language model that functions as an internal adversary to test the resilience of its systems against cyber threats. The model was employed during the training of the latest GPT‑5.6 release, and the company says it helped make this version its most robust yet.
Summary written by Kernelia from the original article by MIT Technology Review. The story and its rights belong to its author.

