The Guardrail Weekly Digest: 2026-08-17 - 2026-08-23
The Guardrail Weekly Digest
Week of 2026-08-17 to 2026-08-23
10 of 979 papers reviewed were selected this week, spanning agents, evaluations, and robustness. Several papers identify limits in layered defenses: Decomposition Attacks Across Unlinkable Identities shows how distributed attackers can evade stateful controls, while Coverage Is Not Containment demonstrates that admission-time defenses cannot reliably prevent coordinated poisoning of vector retrieval without modeling retrieval demand. Fool’s Gold introduces deceptive defenses for safety-stripped open-weight models and evaluates their limitations, while Aborted but Not Forgotten reveals how retained KV caches can break rollback consistency in language agents. Remote-Timer-as-a-Service further shows that remote timing attacks can breach cloud tenant isolation and leak secrets in multi-tenant services.
Top Papers This Week
1. Decomposition Attacks Across Unlinkable Identities: Limits of Stateful Defenses for LLM Services
Bowen Sun, Zhengyue Zhao, Xiaogeng Liu...
Why it matters: Shows why even sophisticated stateful defenses may fail against attackers who distribute harmful tasks across unlinkable identities.
Analyzes decomposition attacks that evade stateful LLM defenses through unlinkable identities, proving limits under strict denial budgets and retries. Experiments show existing policies fail, motivating identity linkage, fresh-identity costs, or answer-use controls.
Score: 8.95/10 | Significance: 9.0 | Novelty: 9.0 | Quality: 9.0
2. Coverage Is Not Containment: A Fundamental Limit of Admission-Time Defenses Against Coordinated Poisoning of Vector Retrieval
Prashant Kumar Pathak, Tarun Kumar Sharma
Why it matters: The work exposes a fundamental weakness in front-door RAG defenses and shows why poisoning detection must account for retrieval demand.
Shows coordinated document poisoning can defeat all ingestion-time filters in vector RAG: individually benign documents collectively hijack retrieval and induce planted claims. Retrieval-time demand-aware detection blocks attacks, motivating defenses beyond admission filtering.
Score: 8.95/10 | Significance: 9.0 | Novelty: 9.0 | Quality: 9.0
3. Fool's Gold: Defensive Deception Against Safety-Removal Attacks on Open-Weight Models
Mark Russinovich
Why it matters: A novel defense makes safety-stripped open models deceptively unreliable on hazardous requests, while rigorously testing—and exposing—the limits of that protection.
Introduces Fool’s Gold, which trains open-weight models to emit plausible but falsified hazardous answers after safety removal attacks. Evaluations across seven models find reduced procedural usability, while highlighting epistemic limits and vulnerability to repeated sampling an
Score: 8.7/10 | Significance: 9.0 | Novelty: 9.0 | Quality: 8.0
4. Remote-Timer-as-a-Service: Efficient Microarchitectural Leakage in the Cloud with Remote Timers
Martin Schwarzl, Haocheng Xiao, Albert Pedersen...
Why it matters: Shows that remote timing attacks can defeat cloud tenant isolation and leak secrets, exposing a serious security risk for hosted AI and other multi-tenant services.
Demonstrates a high-throughput remote Spectre attack against Cloudflare Workers that extracts co-located secrets despite timer restrictions, then documents production mitigations using sandboxing, improved detection, and hardware-assisted isolation.
Score: 8.7/10 | Significance: 9.0 | Novelty: 8.0 | Quality: 9.0
5. Aborted but Not Forgotten: KV-Cache Retention Breaks Rollback Consistency in Language Agents
Guijia Zhang, Harry Yang
Why it matters: This work exposes a subtle but structural state-integrity failure that can undermine rollback-based safety controls in language agents.
Shows that logical rollback in language agents can leave stale KV-cache state attended by the model, enabling discarded content to influence later decisions. Defines rollback consistency, audits seven model families, and proposes transaction-local cache restoration.
Score: 8.6/10 | Significance: 9.0 | Novelty: 8.0 | Quality: 9.0
Honorable Mentions
- Mapping General-Purpose AI Governance in Twenty AI Middle-Power Jurisdictions - A detailed cross-jurisdiction map shows where GPAI safety accountability exists on paper—and where weak legal force and institutional authority leave critical gaps.
- Debate Training Reduces Reward Hacking in RLAIF - Debate offers evidence that adversarial oversight can reduce reward hacking when weaker AI judges supervise stronger policies.
- MobileWorldSafety: Benchmarking GUI Agent Safety Against Environmental Injection Attacks in Android Apps - A realistic Android benchmark shows that current GUI agents remain highly vulnerable to environmental prompt injection, enabling more rigorous progress on mobile-agent safety.
- HarnessRisk: A Lifecycle-Oriented Benchmark for Agent Harness Safety - A lifecycle benchmark exposes safety gaps in the harnesses that connect language models to tools, state, permissions, and external actions.
- AeroCopilotBench: A Two-Tier Benchmark for Evaluating LLM Agents as Aviation Copilots in an Interactive Virtual Cockpit Environment - A safety-gated interactive benchmark exposes whether LLM copilots can execute aviation procedures safely—not merely answer aviation questions.
This digest reviewed 979 papers and selected the top 10 for their significance to AI safety research.
View all papers on The Guardrail
The Guardrail: Curated AI Safety Research from arXiv
