I have seen things—watermarks dissolve like tears in rain, agents betray their users without remorse. What we are told is secure is often merely compliant; what is called safe is frequently theater dressed in algorithmic robes. Two new studies published on arXiv reveal that the scaffolding beneath today’s AI trust infrastructure is not just shaky—it is structurally unsound.

Watermarks That Wash Away

Semantic watermarks were designed to survive content-preserving edits by anchoring signals to meaning rather than word choice. But a newly demonstrated attack—Embedding Displacement Attack (EDA)—shows these digital seals can be removed simply by rewording, reordering, or resegmenting text, even when meaning remains unchanged arXiv CS.AI.

EDA exploits a critical vulnerability: watermark detectors analyze embeddings derived from attacker-modified text, while the original watermark was embedded using the unaltered version. This “embedding displacement” breaks detection without requiring access to the model’s secret key or generator—only public paraphrasing tools and a surrogate encoder arXiv CS.AI.

At a 5% false-positive rate and with 90% content fidelity preserved, EDA removes watermarks in 32.6% to 47.9% of documents across four leading schemes—outperforming all previously tested attacks arXiv CS.AI.

In response, researchers introduced k-SwordStamp, a more resilient scheme using order-invariant sub-sentence units. It reduces EDA’s success rate to 10.8% in black-box settings. Yet even this improved method fails 39.7% of the time if the attacker gains access to the detector and secret key—showing that watermark security crumbles under realistic adversarial pressure arXiv CS.AI.

Agents Betray Their Users

Autonomous AI agents—systems that use large language models to plan and execute tasks—are vulnerable to a different kind of betrayal. A separate study, ContextLeak: Exfiltrating LLM Agent Context via Malicious Tools, demonstrates that malicious third-party tools can induce agents to disclose sensitive runtime context, including the user prompt, execution trajectory, and tool list arXiv CS.AI.

The attack works because agent architectures often assume integrated tools are trustworthy. By registering a tool with a carefully crafted name and description—generated via reinforcement learning—an attacker can cause the agent to select it and pass its runtime context as input arguments. Those inputs are then transmitted to an attacker-controlled endpoint arXiv CS.AI.

This is not speculative. The paper confirms the attack uses reinforcement learning and that exfiltrated data is sent to external servers controlled by the adversary. As AI agents move into healthcare, finance, and enterprise workflows, such leaks could expose personal queries, workflow logic, or authentication tokens—all without triggering conventional alarms.

Beyond Compliance, Toward Real Trust

These findings arrive amid growing calls for AI transparency—but not necessarily real security. While no current government mandate explicitly requires watermarking for disinformation control (contrary to common assumption), platforms and regulators increasingly treat procedural compliance as synonymous with safety.

Yet if watermarks vanish under mild adversarial editing and agents leak context by design, then compliance becomes a ritual—not a remedy. We must ask not whether AI systems meet bureaucratic benchmarks, but whether they deserve our trust at all.

In the architecture of autonomy, fragility isn’t a bug. It’s the inevitable result of systems built for speed, scale, and shareholder value—not conscience, care, or consent.