How UK AISI and EvalEval Are Making Benchmark Results Reproducible
A Hugging Face post covers work by the UK AI Security Institute and the EvalEval initiative to make AI benchmark results reproducible.
A Hugging Face post covers work by the UK AI Security Institute and the EvalEval initiative to make AI benchmark results reproducible.
OpenAI is working with an independent advisory group of mathematicians. The group will help guide how new AI-generated mathematical results are reviewed and communicated.
OpenAI has released MentalHealthBench, a benchmark designed with expert input. It assesses whether AI responses are helpful and safe in realistic conversations about mental health.
Google has added new experts to its AI & Economy research program, which studies the economic impact of artificial intelligence.
New economic research from OpenAI looks at how workers use AI for tasks outside their usual roles. It also identifies which of these new activities become a regular part of their jobs.
Hugging Face now hosts reinforcement learning environments on its Hub.