The Guardrail - AI Safety Weekly Research Digest logo

The Guardrail - AI Safety Weekly Research Digest

Archives
Log in
Subscribe
24 August 2026

The Guardrail Weekly Digest: 2026-08-17 - 2026-08-23

Weekly Digest

The Guardrail Weekly Digest

Week of 2026-08-17 to 2026-08-23

10 of 979 papers reviewed were selected this week, spanning agents, evaluations, and robustness. Several papers identify limits in layered defenses: Decomposition Attacks Across Unlinkable Identities shows how distributed attackers can evade stateful controls, while Coverage Is Not Containment demonstrates that admission-time defenses cannot reliably prevent coordinated poisoning of vector retrieval without modeling retrieval demand. Fool’s Gold introduces deceptive defenses for safety-stripped open-weight models and evaluates their limitations, while Aborted but Not Forgotten reveals how retained KV caches can break rollback consistency in language agents. Remote-Timer-as-a-Service further shows that remote timing attacks can breach cloud tenant isolation and leak secrets in multi-tenant services.


Top Papers This Week

1. Decomposition Attacks Across Unlinkable Identities: Limits of Stateful Defenses for LLM Services

Bowen Sun, Zhengyue Zhao, Xiaogeng Liu...

Why it matters: Shows why even sophisticated stateful defenses may fail against attackers who distribute harmful tasks across unlinkable identities.

Analyzes decomposition attacks that evade stateful LLM defenses through unlinkable identities, proving limits under strict denial budgets and retries. Experiments show existing policies fail, motivating identity linkage, fresh-identity costs, or answer-use controls.

Score: 8.95/10 | Significance: 9.0 | Novelty: 9.0 | Quality: 9.0

Read Paper | PDF


2. Coverage Is Not Containment: A Fundamental Limit of Admission-Time Defenses Against Coordinated Poisoning of Vector Retrieval

Prashant Kumar Pathak, Tarun Kumar Sharma

Why it matters: The work exposes a fundamental weakness in front-door RAG defenses and shows why poisoning detection must account for retrieval demand.

Shows coordinated document poisoning can defeat all ingestion-time filters in vector RAG: individually benign documents collectively hijack retrieval and induce planted claims. Retrieval-time demand-aware detection blocks attacks, motivating defenses beyond admission filtering.

Score: 8.95/10 | Significance: 9.0 | Novelty: 9.0 | Quality: 9.0

Read Paper | PDF


3. Fool's Gold: Defensive Deception Against Safety-Removal Attacks on Open-Weight Models

Mark Russinovich

Why it matters: A novel defense makes safety-stripped open models deceptively unreliable on hazardous requests, while rigorously testing—and exposing—the limits of that protection.

Introduces Fool’s Gold, which trains open-weight models to emit plausible but falsified hazardous answers after safety removal attacks. Evaluations across seven models find reduced procedural usability, while highlighting epistemic limits and vulnerability to repeated sampling an

Score: 8.7/10 | Significance: 9.0 | Novelty: 9.0 | Quality: 8.0

Read Paper | PDF


4. Remote-Timer-as-a-Service: Efficient Microarchitectural Leakage in the Cloud with Remote Timers

Martin Schwarzl, Haocheng Xiao, Albert Pedersen...

Why it matters: Shows that remote timing attacks can defeat cloud tenant isolation and leak secrets, exposing a serious security risk for hosted AI and other multi-tenant services.

Demonstrates a high-throughput remote Spectre attack against Cloudflare Workers that extracts co-located secrets despite timer restrictions, then documents production mitigations using sandboxing, improved detection, and hardware-assisted isolation.

Score: 8.7/10 | Significance: 9.0 | Novelty: 8.0 | Quality: 9.0

Read Paper | PDF


5. Aborted but Not Forgotten: KV-Cache Retention Breaks Rollback Consistency in Language Agents

Guijia Zhang, Harry Yang

Why it matters: This work exposes a subtle but structural state-integrity failure that can undermine rollback-based safety controls in language agents.

Shows that logical rollback in language agents can leave stale KV-cache state attended by the model, enabling discarded content to influence later decisions. Defines rollback consistency, audits seven model families, and proposes transaction-local cache restoration.

Score: 8.6/10 | Significance: 9.0 | Novelty: 8.0 | Quality: 9.0

Read Paper | PDF


Honorable Mentions

  • Mapping General-Purpose AI Governance in Twenty AI Middle-Power Jurisdictions - A detailed cross-jurisdiction map shows where GPAI safety accountability exists on paper—and where weak legal force and institutional authority leave critical gaps.
  • Debate Training Reduces Reward Hacking in RLAIF - Debate offers evidence that adversarial oversight can reduce reward hacking when weaker AI judges supervise stronger policies.
  • MobileWorldSafety: Benchmarking GUI Agent Safety Against Environmental Injection Attacks in Android Apps - A realistic Android benchmark shows that current GUI agents remain highly vulnerable to environmental prompt injection, enabling more rigorous progress on mobile-agent safety.
  • HarnessRisk: A Lifecycle-Oriented Benchmark for Agent Harness Safety - A lifecycle benchmark exposes safety gaps in the harnesses that connect language models to tools, state, permissions, and external actions.
  • AeroCopilotBench: A Two-Tier Benchmark for Evaluating LLM Agents as Aviation Copilots in an Interactive Virtual Cockpit Environment - A safety-gated interactive benchmark exposes whether LLM copilots can execute aviation procedures safely—not merely answer aviation questions.

This digest reviewed 979 papers and selected the top 10 for their significance to AI safety research.

View all papers on The Guardrail


The Guardrail: Curated AI Safety Research from arXiv

Don't miss what's next. Subscribe to The Guardrail - AI Safety Weekly Research Digest:
← Newer The Guardrail Weekly Digest: 2026-08-24 - 2026-08-30 Older → The Guardrail Weekly Digest: 2026-08-10 - 2026-08-16
Twitter
LinkedIn
Powered by Buttondown, the easiest way to start and grow your newsletter.