LLM Watermarking Can Alter AI Agent Tool Calls and Behaviour: Study
• Lasso Security conducted a study finding that LLM watermarking, a method used to identify AI-generated text, alters how open-weight models call tools and behave. • The researchers observed that these watermarking techniques changed individual tool-call outcomes and refusal behaviors, with the most pronounced effects occurring during prompt injection attacks. • These findings suggest that security measures intended to track AI content can inadvertently degrade the reliability and functional accuracy of AI agents.
analyticsindiamag.com
![[AINews] GPT-6 Astra: OpenAI’s biggest LLM launch of all time](/_next/image?url=https%3A%2F%2Fsubstackcdn.com%2Fimage%2Ffetch%2F%24s_!75mH!%2Cw_1200%2Ch_675%2Cc_fill%2Cf_jpg%2Cq_auto%3Agood%2Cfl_progressive%3Asteep%2Cg_auto%2Fhttps%253A%252F%252Fsubstack-post-media.s3.amazonaws.com%252Fpublic%252Fimages%252F3e58156f-49e2-48e8-af49-ce5edd8e68b6_1118x1118.png&w=1920&q=85)
