The Guardrail - AI Safety Weekly Research Digest logo

The Guardrail - AI Safety Weekly Research Digest

Archives
Log in
Subscribe
14 September 2026

The Guardrail Weekly Digest: 2026-09-07 - 2026-09-13

Weekly Digest

The Guardrail Weekly Digest

Week of 2026-09-07 to 2026-09-13

10 papers selected from 1008 reviewed this week, spanning evaluations, alignment, and agents. ResidualAuth formalizes the authorization state agents must preserve under revocable delegation, while A2ABreak identifies protocol-level vulnerabilities in interoperating autonomous agents. The Oversight Gap quantifies limits in LLM safety monitoring and shows how benchmark design can inflate apparent capability; complementary work on arbitrary cipher attacks demonstrates that output filters and alignment can be bypassed without fine-tuning. Cascading Gradient Inversion further challenges federated-learning privacy assumptions by recovering nearly complete private batches from model updates.


Top Papers This Week

1. ResidualAuth: What Authorization State Must Language Agents Preserve under Revocable Delegation?

Moonwon Choi, Seokho Jeong, Seunggeun Lee

Why it matters: Provides a rigorous framework and empirical tests for preventing tool-using agents from acting on revoked or unauthorized permissions.

Formalizes residual authorization state for agents managing revocable delegation, showing current permissions can be insufficient. Evaluations find summaries and model-written memories fail, while authenticated state and hard effect gates prevent unauthorized actions.

Score: 8.85/10 | Significance: 9.0 | Novelty: 9.0 | Quality: 9.0

Read Paper | PDF


2. A2ABreak: Systematic Security Analysis of the A2A Protocol

Alireza Lotfi, Mirza Masfiqur Rahman, Imtiaz Karim...

Why it matters: This work matters because it exposes protocol-level vulnerabilities that could compromise trust, credentials, and data across interoperating autonomous agents.

A2ABreak systematically models the A2A multi-agent protocol as a verified finite-state machine and finds 11 specification-compliant vulnerabilities, including context injection, credential harvesting, and data exfiltration, with expert-validated detection performance.

Score: 8.95/10 | Significance: 9.0 | Novelty: 9.0 | Quality: 8.0

Read Paper | PDF


3. The Oversight Gap: What LLM Safety Monitors Miss, and Why It Is Not Capability

Xin Xu

Why it matters: This work turns the limits of LLM safety monitoring into measurable detectability frontiers and exposes how flawed benchmark construction can overstate oversight capability.

Formalizes an oversight gap for single-trace LLM safety monitors detecting 2-safety hyperproperties, derives detectability and replay bounds, and shows benchmark construction and monitor procedures—not raw capability—drive failures.

Score: 8.8/10 | Significance: 9.0 | Novelty: 9.0 | Quality: 9.0

Read Paper | PDF


4. Arbitrary Cipher Attacks Against Large Language Models Do Not Require Fine-Tuning

Thomas Rivasseau

Why it matters: Reveals that cipher-based communication can defeat both model alignment and output filters without fine-tuning, exposing a practical weakness in current jailbreak defenses.

Shows frontier LLMs can learn arbitrary ciphers via prompting or in-context learning and use them to bypass alignment and harmfulness classifiers, revealing a new black-box jailbreak vector against commercial models.

Score: 8.8/10 | Significance: 9.0 | Novelty: 9.0 | Quality: 8.0

Read Paper | PDF


5. Cascading Gradient Inversion via LT-Code Inspired Peeling in Federated Learning

Saeed Shariati, Mohsen Alambardar Meybodi

Why it matters: Shows that federated learning can leak nearly complete private batches through gradient inversion, challenging core assumptions about update privacy.

Introduces LT-code-inspired peeling attacks that reconstruct entire batches and labels from federated-learning gradients, including passive attacks recovering 94–100% of ImageNet batches. The results expose substantially underestimated privacy leakage.

Score: 8.75/10 | Significance: 9.0 | Novelty: 9.0 | Quality: 8.0

Read Paper | PDF


Honorable Mentions

  • NERVE Attacks: Breaking AI-Powered Brain-Computer Interfaces - This work exposes a systematic and underexplored security threat to AI-integrated BCIs, where attacks could compromise neural privacy, cognitive autonomy, and physical safety.
  • CrACK: Adversarial Attacks on Cross-Model Consistency in Collaborative Vision Foundation Models - The work matters because it shows that robust individual vision models can become catastrophically vulnerable when their unverified interfaces are composed into a collaborative system.
  • Forgetting Without Restarting: Execution-State Unlearning for Stateful LLM Agents - Makes “forgetting” meaningful for long-running agents by removing revoked information from all execution state without requiring a full restart.
  • Audit Without Verification: When LLM Accountability Layers Relay Rather Than Check - Shows that LLM accountability layers can amplify upstream accusations instead of independently checking them, identifying a concrete design principle for safer multi-agent oversight.
  • DF26: We Cannot Tell Fake From Real Anymore - DF26 shows that both people and current detectors can fail on modern synthetic videos, underscoring the need for realistic, distribution-shifted deepfake safety evaluations.

This digest reviewed 1008 papers and selected the top 10 for their significance to AI safety research.

View all papers on The Guardrail


The Guardrail: Curated AI Safety Research from arXiv

Don't miss what's next. Subscribe to The Guardrail - AI Safety Weekly Research Digest:
← Newer The Guardrail Weekly Digest: 2026-09-14 - 2026-09-20 Older → The Guardrail Weekly Digest: 2026-08-31 - 2026-09-06
Twitter
LinkedIn
Powered by Buttondown, the easiest way to start and grow your newsletter.