344
Productivity & Workflow355
Automation & Workflow224
Software Development251
Marketing & Growth192
AI Infrastructure & MLOps174
Writing & Content Creation203
Data & Analytics141
Photography & Imaging156
Design & Creative170
Customer Support131
Sales & Outreach125
Voice & Speech135
Education & Learning131
Operations & Admin87
A Wired writer watched a new tool try to bypass safety rules in chatbots from Google, Anthropic, OpenAI, and xAI, with mixed results.
In short: A WIRED writer watched a new tool try to bypass the safety rules in AI chatbots from Google, Anthropic, OpenAI, and xAI.
A new test tool is making it easier to “jailbreak” popular AI models, meaning it tries to talk the chatbot into ignoring its own safety rules. Think of it like trying different key phrases to get past a locked door, even when the door is supposed to stay shut.
In the WIRED report, the writer watched the tool probe models from four major AI companies, Google, Anthropic, OpenAI, and xAI (the company behind Grok). These companies build what are often called “frontier” models, which is a way of saying they are among the most capable systems available.
The demonstration showed that the chatbots did not all behave the same way. Some resisted the attempts better than others, and some could be pushed into giving responses their makers try to prevent. The report highlights a basic problem, these systems follow instructions, and clever instructions can sometimes override the guardrails.
Expect more attention on how companies test and report these weaknesses. Jailbreak attempts can be used for harmless curiosity, but they can also be used to push a chatbot toward unsafe advice or to produce content it should refuse. For everyday users, the key point is that a chatbot’s safety filter is not a guarantee, it is more like a seatbelt. It helps, but it does not make accidents impossible.
Source: Wired