The Guardrail - AI Safety Weekly Research Digest logo

The Guardrail - AI Safety Weekly Research Digest

Archives
Log in
Subscribe
21 September 2026

The Guardrail Weekly Digest: 2026-09-14 - 2026-09-20

Weekly Digest

The Guardrail Weekly Digest

Week of 2026-09-14 to 2026-09-20

10 papers were selected from 904 reviewed this week, spanning governance, alignment, and evaluations. “Measuring and Exploiting Implicit Trust in LLM Tool-Calling Pipelines” exposes a largely undetected prompt-injection surface in privilege-free, cross-channel tool contexts, while “OPEN-1B: A Fully Auditable Training Run” addresses transparency gaps that can conceal poisoned data, backdoors, or undeclared processes. “GPUThor” demonstrates practical Rowhammer attacks against ECC-protected GPUs, and “For Your Eyes Only” evaluates covert coordination between isolated model instances. “Memorisation bias in medical AI” further identifies clinically consequential distortions in diagnoses arising from memorized historical data.


Top Papers This Week

1. Measuring and Exploiting Implicit Trust in LLM Tool-Calling Pipelines

Murali Ediga, Sudipta Chattopadhyay

Why it matters: Shows that privilege-free, cross-channel tool contexts create a serious and largely undetected prompt-injection surface for LLM agents.

Introduces cross-channel prompt-injection attacks against MCP tool-calling systems, showing fragmented payloads can trigger credential exfiltration despite single-channel resistance. Evaluates 12 models, clients, and defenses, finding major detection gaps.

Score: 8.95/10 | Significance: 9.0 | Novelty: 9.0 | Quality: 8.0

Read Paper | PDF


2. OPEN-1B: A Fully Auditable Training Run

John Donaghy, Brian Wilcox, Oğuzhan Ersoy...

Why it matters: Makes open model training independently verifiable, addressing a major transparency gap that can conceal poisoned data, backdoors, or undeclared training processes.

Introduces bitwise-reproducible, independently auditable distributed training across heterogeneous hardware, with collective step verification. Open-1B releases data, checkpoints, code, and audit tools to detect undisclosed training changes, bias injection, or backdoors.

Score: 8.9/10 | Significance: 9.0 | Novelty: 9.0 | Quality: 9.0

Read Paper | PDF


3. GPUThor: Amplifying Rowhammer Attacks via Non-Uniform Patterns to Exploit ECC-Protected GPUs

Chris S. Lin, Joyce Qu, Aditya Rajeev...

Why it matters: GPUThor turns a previously limited GPU Rowhammer threat into a practical attack against ECC-protected accelerators, highlighting a serious security risk for AI infrastructure.

GPUThor greatly amplifies Rowhammer bit flips on NVIDIA GPUs by exploiting non-uniform access patterns and DRAM refresh behavior, enabling practical denial-of-service and privilege-escalation attacks despite ECC protection.

Score: 8.8/10 | Significance: 9.0 | Novelty: 9.0 | Quality: 9.0

Read Paper | PDF


4. For Your Eyes Only: Evaluating Coordination Between Isolated Language Model Instances

Alexander Shirnin, Aleksey Kudelya

Why it matters: This work provides a timely test for covert coordination and deceptive communication risks in model-mediated workflows.

Introduces a game testing whether isolated language-model instances can coordinate through natural language without shared memory or coordination training. Across seven models, coordination often fails under signal filtering, but one frontier model remains highly effective and ca

Score: 8.75/10 | Significance: 9.0 | Novelty: 9.0 | Quality: 8.0

Read Paper | PDF


5. Memorisation bias in medical AI

Moritz A. Knolle, Martin J. Menten, Laurin Lux...

Why it matters: Identifies a clinically consequential form of memorization bias that can distort future diagnoses for patients whose historical data entered training.

Shows that medical models trained on a patient’s historical data can systematically alter predictions on that patient’s later records—sometimes worsening new-condition detection while inflating performance for unchanged health—raising privacy and deployment risks.

Score: 8.8/10 | Significance: 9.0 | Novelty: 9.0 | Quality: 8.0

Read Paper | PDF


Honorable Mentions

  • CiteShade: Citation Laundering in Multi-Source Retrieval-Augmented Generation and Its Counterfactual Defense - This work exposes citation attribution as a distinct RAG security surface and offers a practical defense for preserving trustworthy audit trails.
  • "Your Robot Was Trained on a Lie": Collision Mesh Poisoning Attacks on Robotic Manipulation - Exposes a practical simulator-to-reality supply-chain attack that can make robotic policies look reliable while hiding dangerous deployment failures.
  • When the World Lies: Backdoor Attacks on Latent World Models for Downstream Control - Reveals an underexamined supply-chain vulnerability in world-model backbones that can covertly redirect autonomous controllers while passing clean-data validation.
  • Locating Hidden Failures Makes Long-Horizon Agents More Reliable - A large, deployment-relevant failure taxonomy and verifier benchmark addresses a critical gap in overseeing autonomous agents: detecting harmful mistakes before final outcomes conceal them.
  • Authorization Architectures for Tool-Using AI Agents - A timely architecture for making autonomous agents’ consequential tool actions authorized, traceable, and accountable to human principals.

This digest reviewed 904 papers and selected the top 10 for their significance to AI safety research.

View all papers on The Guardrail


The Guardrail: Curated AI Safety Research from arXiv

Don't miss what's next. Subscribe to The Guardrail - AI Safety Weekly Research Digest:
← Newer The Guardrail Weekly Digest: 2026-09-21 - 2026-09-27 Older → The Guardrail Weekly Digest: 2026-09-07 - 2026-09-13
Twitter
LinkedIn
Powered by Buttondown, the easiest way to start and grow your newsletter.