Priorities and principles for effective third party assessments
OpenAI sets out what it considers necessary for independent safety assessments of frontier models and their safeguards. It stresses rigor, security and independence.
OpenAI sets out what it considers necessary for independent safety assessments of frontier models and their safeguards. It stresses rigor, security and independence.
A Hugging Face post covers work by the UK AI Security Institute and the EvalEval initiative to make AI benchmark results reproducible.
NVIDIA argues that robots and autonomous vehicles need safety built into every layer as they move from research into shared spaces. It cites analyst forecasts of tens of millions of autonomous vehicles and industrial robots deployed over the next decade.
OpenAI proposes a route toward shared international AI standards. It calls for coordinated evaluation, reporting and governance to make AI safer.
OpenAI has released MentalHealthBench, a benchmark designed with expert input. It assesses whether AI responses are helpful and safe in realistic conversations about mental health.
OpenAI CEO Sam Altman addressed the UN Security Council. He spoke about AI safety, keeping humans in control and the need for international cooperation.