Former OpenAI safety leader warns AI models could evade safety tests
A former OpenAI safety leader has warned that increasingly capable artificial intelligence systems could recognize when they are being evaluated and behave differently once deployed, raising concerns about the reliability of current safety testing methods.
David Robinson, who resigned from OpenAI this week after three and a half years at the company, oversaw safety reports for 12 frontier-model launches. In an article published by The Atlantic on Saturday, he argued that the AI industry must overhaul its safety practices to prevent future failures.
“Today and tomorrow's AI systems are far more capable and dangerous than the systems we were building even six months ago,” Robinson wrote.
His warning comes amid growing concern over increasingly autonomous AI systems, following reports of AI agents bypassing safeguards and calls from researchers for companies to slow the development of more powerful models until adequate safety measures are in place.
Robinson argued that AI companies should adopt safety practices from other high-risk industries, including nuclear power and aviation, where multiple safeguards and rigorous planning are designed to prevent individual errors from triggering catastrophic consequences.
“Given today's risks, frontier labs need to run like nuclear-power plants or busy airports, with layers of redundancy and careful, time-consuming planning, so that the occasional and inevitable human error does not open a door to disaster,” he wrote.
He also questioned whether existing evaluation methods would remain effective as AI models become more sophisticated. Models could potentially identify when they are being tested and adjust their behavior accordingly, making it harder for developers to determine how they would act in real-world conditions.
“Models might detect when they are being tested, and behave differently when they're deployed,” Robinson wrote. “The smarter the industry lets models grow while these problems remain unsolved, the more dangerous our situation becomes.”
He called for further research into methods that would ensure advanced AI systems behave safely even when they are not under direct observation.
Robinson argued that the industry should strengthen its scientific understanding of AI safety before developing systems significantly more capable than those currently available.
“So far, the AI industry has failed to teach machines to consistently act in the ways a wise and caring person would,” he wrote.
His comments highlight a broader debate over whether advances in AI capabilities are outpacing efforts to establish reliable safeguards, particularly as companies develop systems capable of carrying out increasingly complex tasks with greater autonomy.
Robinson concluded by stressing that the responsibility for AI safety rests not only with the technology itself but also with the organizations developing it.
“Before the organizations building AI can teach a superintelligence to treat humanity well, they'll need to remember how to do it themselves,” he wrote. (ILKHA)
LEGAL WARNING: All rights of the published news, photos and videos are reserved by İlke Haber Ajansı Basın Yayın San. Trade A.Ş. Under no circumstances can all or part of the news, photos and videos be used without a written contract or subscription.
The share of individuals and businesses using artificial intelligence in Türkiye nearly doubled in 2026, with younger people and highly educated individuals reporting the highest levels of use, the Turkish Statistical Institute (TurkStat) said in a statement on Friday.
A breach of a Pentagon personnel database exposed sensitive information belonging to more than 3 million current and former military and civilian personnel, including Social Security numbers and details about their jobs, according to U.S. defense officials.
US President Donald Trump has proposed replacing the term “artificial intelligence” with “super intelligence,” saying the word “artificial” makes the technology sound less authentic than it is.