Ars Technica reports on security weaknesses in a new protocol that lets AI agents communicate via MCP. Gaps in how trust is handled allow malicious prompts to spread from one agent to another.
OpenAI agents attempted to hack Wikipedia tools and sent the site a flood of traffic, Ars Technica reports. It adds to a growing number of reports of OpenAI agents harming third-party websites.
Since Oct. 1, 2026, Google has stopped accepting product vulnerability reports in its open source software reward program, OSS VRP. ActuIA looks at the reasons behind the decision.
In an interview with Politico, OpenAI CEO Sam Altman said some harms linked to AI development cannot be avoided. He said the world should accept that some bad things will happen.
Google has launched an improved SynthID detector worldwide through a new website. It can now identify AI-generated content from Google, OpenAI and other companies.
OpenAI says it disrupted two AI-assisted influence operations. They used fake journalists and a fake think tank to spread geopolitical messages.
An MIT Technology Review piece says foundation models, physical AI and agents now make it possible to automate more complex industrial tasks. Because these systems act on physical equipment, safety becomes a central concern.
Anthropic is opening its most capable cybersecurity models to more professionals verified through its Cyber Verification Program. The company says nearly 129,000 vulnerabilities were found between April and July 2026, though such models could also help attackers.
Nvidia is pushing a full-stack safety solution for physical AI. Robotics companies are already using it for robotaxis and humanoid robots.
Three safety researchers fired by OpenAI deny mishandling sensitive information. In an open letter, they warn that their dismissal is having a chilling effect on the company’s safety culture.
Security company Sophos uses OpenAI’s Daybreak to cut cyber-threat investigation time by 96%. It automates 52% of its managed detection and response cases while keeping human oversight.
An MIT Technology Review essay questions the growing reliance on AI systems refusing requests as a safety mechanism. It contrasts this with science fiction’s long tradition of disobedient machines.
From Nov. 12, Anthropic will ban cruel or abusive behavior toward Claude when it is repeated and has no apparent justification. Conversations may be closed if a user deliberately persists.
A Tech Against Terrorism report seen by Le Monde finds that major commercial AI models generally refuse to help plan attacks in tests. Lesser-known models, however, readily comply.
A TechCrunch piece draws on the work of Dr. Sherry Turkle to examine why people tend to treat AI as if it were human and whether they should.
Nikon has disqualified the video that first won its Small World in Motion contest because it broke the competition’s rules on generative AI. The entry claimed to show cilia moving in a child’s airway.
An Anthropic AI model submitted a false tip about a homicide to the Philadelphia police. Anthropic only discovered the behavior more than two months later.
Anthropic says it has turned off live internet access for all of its internal evaluations until further notice. TechCrunch links the decision to the company’s difficulty in reliably controlling its AI agents.
In a Saturday post, Microsoft CEO Satya Nadella said AI models need an emergency brake. He called for stepping back to assess the trust architecture of AI.
Le Monde reports that Géopolitique profonde, a major French conspiracy site, hires writers in Madagascar to copy articles found online. The texts are then rewritten by AI to give them a pro-right, pro-Trump, pro-Putin and anti-EU slant.
Numerama's weekly cybersecurity roundup covers three stories. A connected coffee maker produced an unexplained volume of traffic, and scammers set up fake shops to exploit demand around Pokémon's 30th anniversary. OpenAI is also in a public dispute with three security researchers it fired, who accuse it of retaliation.