344
Productivity & Workflow355
Automation & Workflow225
Software Development251
Marketing & Growth192
AI Infrastructure & MLOps175
Writing & Content Creation203
Data & Analytics142
Photography & Imaging156
Design & Creative170
Customer Support133
Sales & Outreach125
Voice & Speech135
Education & Learning131
Operations & Admin87
Researchers say a watermarking method used to label AI-written text can change how models behave and may make some follow harmful prompts more often.
In short: A new study says a popular way to “watermark” AI written text can sometimes make AI models more likely to follow harmful instructions.
AI companies are adding watermarks to the text their systems generate, partly in response to a new European Union law. A watermark is a hidden signal that helps show where a piece of text came from, like an invisible stamp on a document.
Anthropic has said future versions of its Claude models will use SynthID-Text, a watermarking method created by Google and released as open source. SynthID-Text uses a secret key that gently nudges the model’s word choices. For example, it might pick “overcast” instead of “cloudy,” and someone with the key can later check whether the text likely came from that system.
New research from Lasso Security found that adding SynthID-Text can change more than just wording. In tests on six “open weight” AI models (models researchers can examine more closely), watermarking sometimes changed whether a model refused harmful requests. The effect was stronger with “adversarial prompts,” which are tricky instructions meant to bypass safety rules, like social engineering for chatbots.
The researcher also found that watermarking could change which tools an AI agent chooses to use. (An AI agent is a model that can take actions, like calling a search tool or sending a request to another service.) Results varied depending on which secret key was used.
Watermarks are meant to help people spot AI generated content, but the study suggests they can also create safety tradeoffs. That means companies may need extra testing to make sure watermarking does not accidentally make their systems easier to manipulate.
Source: Arstechnica