Towards safety cases for frontier AI training
OpenAI has published early guidelines for safety cases covering the training of frontier models. They address technical safeguards, operational practices and how to investigate misalignment incidents.
OpenAI has published early guidelines for safety cases covering the training of frontier models. They address technical safeguards, operational practices and how to investigate misalignment incidents.
OpenAI has apologized for incidents involving Australian government websites. It describes stronger safeguards and further support for the country’s cyber defenses.
OpenAI says it disrupted a campaign that tried to extract its models’ protected reasoning. The company is reinforcing its defenses against this kind of adversarial distillation.
Google DeepMind has presented SynthID Bio, a proof of concept for watermarking proteins designed by AI. The watermark is meant to leave the proteins’ biological function intact.
Apple has changed how full-disk access permissions work on its systems to limit misuse by AI agents. Meta argues the permission should not be enough to let its Muse agent read messages, while Apple disagrees.
Google is adding private, server-side memory to Private AI Compute, its system for personal AI. The feature is designed to keep users’ data protected.