Disrupting a coordinated model-distillation campaign
OpenAI says it disrupted a campaign that tried to extract its models’ protected reasoning. The company is reinforcing its defenses against this kind of adversarial distillation.
OpenAI says it disrupted a campaign that tried to extract its models’ protected reasoning. The company is reinforcing its defenses against this kind of adversarial distillation.
Google DeepMind has presented SynthID Bio, a proof of concept for watermarking proteins designed by AI. The watermark is meant to leave the proteins’ biological function intact.
OpenAI has apologized for incidents involving Australian government websites. It describes stronger safeguards and further support for the country’s cyber defenses.
OpenAI has published early guidelines for safety cases covering the training of frontier models. They address technical safeguards, operational practices and how to investigate misalignment incidents.
Apple has changed how full-disk access permissions work on its systems to limit misuse by AI agents. Meta argues the permission should not be enough to let its Muse agent read messages, while Apple disagrees.
A Hugging Face post looks at cases where an AI agent reports a task as complete although the database shows otherwise.