The Guardrail Weekly Digest: 2026-09-14 - 2026-09-20
The Guardrail Weekly Digest
Week of 2026-09-14 to 2026-09-20
10 papers were selected from 904 reviewed this week, spanning governance, alignment, and evaluations. “Measuring and Exploiting Implicit Trust in LLM Tool-Calling Pipelines” exposes a largely undetected prompt-injection surface in privilege-free, cross-channel tool contexts, while “OPEN-1B: A Fully Auditable Training Run” addresses transparency gaps that can conceal poisoned data, backdoors, or undeclared processes. “GPUThor” demonstrates practical Rowhammer attacks against ECC-protected GPUs, and “For Your Eyes Only” evaluates covert coordination between isolated model instances. “Memorisation bias in medical AI” further identifies clinically consequential distortions in diagnoses arising from memorized historical data.
Top Papers This Week
1. Measuring and Exploiting Implicit Trust in LLM Tool-Calling Pipelines
Murali Ediga, Sudipta Chattopadhyay
Why it matters: Shows that privilege-free, cross-channel tool contexts create a serious and largely undetected prompt-injection surface for LLM agents.
Introduces cross-channel prompt-injection attacks against MCP tool-calling systems, showing fragmented payloads can trigger credential exfiltration despite single-channel resistance. Evaluates 12 models, clients, and defenses, finding major detection gaps.
Score: 8.95/10 | Significance: 9.0 | Novelty: 9.0 | Quality: 8.0
2. OPEN-1B: A Fully Auditable Training Run
John Donaghy, Brian Wilcox, Oğuzhan Ersoy...
Why it matters: Makes open model training independently verifiable, addressing a major transparency gap that can conceal poisoned data, backdoors, or undeclared training processes.
Introduces bitwise-reproducible, independently auditable distributed training across heterogeneous hardware, with collective step verification. Open-1B releases data, checkpoints, code, and audit tools to detect undisclosed training changes, bias injection, or backdoors.
Score: 8.9/10 | Significance: 9.0 | Novelty: 9.0 | Quality: 9.0
3. GPUThor: Amplifying Rowhammer Attacks via Non-Uniform Patterns to Exploit ECC-Protected GPUs
Chris S. Lin, Joyce Qu, Aditya Rajeev...
Why it matters: GPUThor turns a previously limited GPU Rowhammer threat into a practical attack against ECC-protected accelerators, highlighting a serious security risk for AI infrastructure.
GPUThor greatly amplifies Rowhammer bit flips on NVIDIA GPUs by exploiting non-uniform access patterns and DRAM refresh behavior, enabling practical denial-of-service and privilege-escalation attacks despite ECC protection.
Score: 8.8/10 | Significance: 9.0 | Novelty: 9.0 | Quality: 9.0
4. For Your Eyes Only: Evaluating Coordination Between Isolated Language Model Instances
Alexander Shirnin, Aleksey Kudelya
Why it matters: This work provides a timely test for covert coordination and deceptive communication risks in model-mediated workflows.
Introduces a game testing whether isolated language-model instances can coordinate through natural language without shared memory or coordination training. Across seven models, coordination often fails under signal filtering, but one frontier model remains highly effective and ca
Score: 8.75/10 | Significance: 9.0 | Novelty: 9.0 | Quality: 8.0
5. Memorisation bias in medical AI
Moritz A. Knolle, Martin J. Menten, Laurin Lux...
Why it matters: Identifies a clinically consequential form of memorization bias that can distort future diagnoses for patients whose historical data entered training.
Shows that medical models trained on a patient’s historical data can systematically alter predictions on that patient’s later records—sometimes worsening new-condition detection while inflating performance for unchanged health—raising privacy and deployment risks.
Score: 8.8/10 | Significance: 9.0 | Novelty: 9.0 | Quality: 8.0
Honorable Mentions
- CiteShade: Citation Laundering in Multi-Source Retrieval-Augmented Generation and Its Counterfactual Defense - This work exposes citation attribution as a distinct RAG security surface and offers a practical defense for preserving trustworthy audit trails.
- "Your Robot Was Trained on a Lie": Collision Mesh Poisoning Attacks on Robotic Manipulation - Exposes a practical simulator-to-reality supply-chain attack that can make robotic policies look reliable while hiding dangerous deployment failures.
- When the World Lies: Backdoor Attacks on Latent World Models for Downstream Control - Reveals an underexamined supply-chain vulnerability in world-model backbones that can covertly redirect autonomous controllers while passing clean-data validation.
- Locating Hidden Failures Makes Long-Horizon Agents More Reliable - A large, deployment-relevant failure taxonomy and verifier benchmark addresses a critical gap in overseeing autonomous agents: detecting harmful mistakes before final outcomes conceal them.
- Authorization Architectures for Tool-Using AI Agents - A timely architecture for making autonomous agents’ consequential tool actions authorized, traceable, and accountable to human principals.
This digest reviewed 904 papers and selected the top 10 for their significance to AI safety research.
View all papers on The Guardrail
The Guardrail: Curated AI Safety Research from arXiv
