SynthID can cause models to follow harmful instructions they would otherwise refuse.
LLMs respond differently to harmful prompts when AI watermarking is used
SynthID can cause models to follow harmful instructions they would otherwise refuse.