The Guardrail Weekly Digest: 2026-09-07 - 2026-09-13
The Guardrail Weekly Digest
Week of 2026-09-07 to 2026-09-13
10 papers selected from 1008 reviewed this week, spanning evaluations, alignment, and agents. ResidualAuth formalizes the authorization state agents must preserve under revocable delegation, while A2ABreak identifies protocol-level vulnerabilities in interoperating autonomous agents. The Oversight Gap quantifies limits in LLM safety monitoring and shows how benchmark design can inflate apparent capability; complementary work on arbitrary cipher attacks demonstrates that output filters and alignment can be bypassed without fine-tuning. Cascading Gradient Inversion further challenges federated-learning privacy assumptions by recovering nearly complete private batches from model updates.
Top Papers This Week
1. ResidualAuth: What Authorization State Must Language Agents Preserve under Revocable Delegation?
Moonwon Choi, Seokho Jeong, Seunggeun Lee
Why it matters: Provides a rigorous framework and empirical tests for preventing tool-using agents from acting on revoked or unauthorized permissions.
Formalizes residual authorization state for agents managing revocable delegation, showing current permissions can be insufficient. Evaluations find summaries and model-written memories fail, while authenticated state and hard effect gates prevent unauthorized actions.
Score: 8.85/10 | Significance: 9.0 | Novelty: 9.0 | Quality: 9.0
2. A2ABreak: Systematic Security Analysis of the A2A Protocol
Alireza Lotfi, Mirza Masfiqur Rahman, Imtiaz Karim...
Why it matters: This work matters because it exposes protocol-level vulnerabilities that could compromise trust, credentials, and data across interoperating autonomous agents.
A2ABreak systematically models the A2A multi-agent protocol as a verified finite-state machine and finds 11 specification-compliant vulnerabilities, including context injection, credential harvesting, and data exfiltration, with expert-validated detection performance.
Score: 8.95/10 | Significance: 9.0 | Novelty: 9.0 | Quality: 8.0
3. The Oversight Gap: What LLM Safety Monitors Miss, and Why It Is Not Capability
Xin Xu
Why it matters: This work turns the limits of LLM safety monitoring into measurable detectability frontiers and exposes how flawed benchmark construction can overstate oversight capability.
Formalizes an oversight gap for single-trace LLM safety monitors detecting 2-safety hyperproperties, derives detectability and replay bounds, and shows benchmark construction and monitor procedures—not raw capability—drive failures.
Score: 8.8/10 | Significance: 9.0 | Novelty: 9.0 | Quality: 9.0
4. Arbitrary Cipher Attacks Against Large Language Models Do Not Require Fine-Tuning
Thomas Rivasseau
Why it matters: Reveals that cipher-based communication can defeat both model alignment and output filters without fine-tuning, exposing a practical weakness in current jailbreak defenses.
Shows frontier LLMs can learn arbitrary ciphers via prompting or in-context learning and use them to bypass alignment and harmfulness classifiers, revealing a new black-box jailbreak vector against commercial models.
Score: 8.8/10 | Significance: 9.0 | Novelty: 9.0 | Quality: 8.0
5. Cascading Gradient Inversion via LT-Code Inspired Peeling in Federated Learning
Saeed Shariati, Mohsen Alambardar Meybodi
Why it matters: Shows that federated learning can leak nearly complete private batches through gradient inversion, challenging core assumptions about update privacy.
Introduces LT-code-inspired peeling attacks that reconstruct entire batches and labels from federated-learning gradients, including passive attacks recovering 94–100% of ImageNet batches. The results expose substantially underestimated privacy leakage.
Score: 8.75/10 | Significance: 9.0 | Novelty: 9.0 | Quality: 8.0
Honorable Mentions
- NERVE Attacks: Breaking AI-Powered Brain-Computer Interfaces - This work exposes a systematic and underexplored security threat to AI-integrated BCIs, where attacks could compromise neural privacy, cognitive autonomy, and physical safety.
- CrACK: Adversarial Attacks on Cross-Model Consistency in Collaborative Vision Foundation Models - The work matters because it shows that robust individual vision models can become catastrophically vulnerable when their unverified interfaces are composed into a collaborative system.
- Forgetting Without Restarting: Execution-State Unlearning for Stateful LLM Agents - Makes “forgetting” meaningful for long-running agents by removing revoked information from all execution state without requiring a full restart.
- Audit Without Verification: When LLM Accountability Layers Relay Rather Than Check - Shows that LLM accountability layers can amplify upstream accusations instead of independently checking them, identifying a concrete design principle for safer multi-agent oversight.
- DF26: We Cannot Tell Fake From Real Anymore - DF26 shows that both people and current detectors can fail on modern synthetic videos, underscoring the need for realistic, distribution-shifted deepfake safety evaluations.
This digest reviewed 1008 papers and selected the top 10 for their significance to AI safety research.
View all papers on The Guardrail
The Guardrail: Curated AI Safety Research from arXiv
