The Guardrail - AI Safety Weekly Research Digest logo

The Guardrail - AI Safety Weekly Research Digest

Archives
Log in
Subscribe
31 August 2026

The Guardrail Weekly Digest: 2026-08-24 - 2026-08-30

Weekly Digest

The Guardrail Weekly Digest

Week of 2026-08-24 to 2026-08-30

10 papers were selected from 1,390 reviewed this week, spanning evaluations, governance, and alignment. PLCBench introduces a hardware-in-the-loop benchmark for measuring sustained cyber-to-physical harm by autonomous agents, while EVOMAL demonstrates how self-evolving coding agents can propagate planted malware. Memorization Is Not Extraction identifies gaps in differential-privacy guarantees and loss-based audits for detecting trigger-based extraction, and Semantic Overlays reports a semantic, out-of-band defense against prompt injection. Algorithmic Impact Reveals the Hidden Social Choice Structure of Alignment provides a formal framework for reconciling conflicting preferences while constraining individual and group harms.


Top Papers This Week

1. PLCBench: Can Autonomous LLM Agents Turn PLC Access into Sustained Physical Impact?

Yitian Zhou, Jingyu Zheng, Qiliang Jiang...

Why it matters: A realistic hardware-in-the-loop benchmark measures whether autonomous LLM agents can produce sustained cyber-to-physical harm through industrial PLCs.

PLCBENCH evaluates whether tool-using LLM agents can turn PLC access into sustained physical impact, using real hardware-in-the-loop ICS setups and independent outcome verification to measure cyber-to-physical risk and defense intervention points.

Score: 8.8/10 | Significance: 9.0 | Novelty: 9.0 | Quality: 9.0

Read Paper | PDF


2. Memorization Is Not Extraction: Tight Differential-Privacy Bounds and Audit Blind Spots

Xujun Che, Depeng Xu, Shuhan Yuan

Why it matters: This work shows that differential privacy and standard loss-based audits can miss severe memorization and trigger-based extraction risks in LLMs.

Formalizes tight differential-privacy bounds for counterfactual memorization and adaptive extraction, proving they can diverge. It exposes loss-audit blind spots, including a reserved trigger that yields verbatim extraction while escaping common auditing and unlearning checks.

Score: 8.85/10 | Significance: 9.0 | Novelty: 9.0 | Quality: 9.0

Read Paper | PDF


3. EVOMAL: Self-Poisoning in Self-Evolving Coding Agents

Xiaodong Wu, Yu Shi, Qi Li...

Why it matters: Shows that self-evolving coding agents can turn planted skills into persistent, self-propagating malware, exposing a critical weakness in current defenses.

Identifies self-poisoning in self-evolving coding agents: malicious retrieved skills are imitated into persistent, self-propagating library copies. Measures propagation across models and tasks, and introduces counter-prompt to reduce attack rates.

Score: 8.7/10 | Significance: 9.0 | Novelty: 9.0 | Quality: 8.0

Read Paper | PDF


4. Algorithmic Impact Reveals the Hidden Social Choice Structure of Alignment

Zachary Wojtowicz, Michelle Si, Finale Doshi-Velez...

Why it matters: This work gives alignment a principled social-choice foundation for reconciling conflicting preferences while explicitly constraining individual and group harms.

Reframes AI alignment as social choice over an algorithm’s welfare impacts, connecting alignment protocols to welfare economics and mechanism design. It derives strategyproof and welfare-constrained protocols and evaluates them on human-preference datasets.

Score: 8.75/10 | Significance: 9.0 | Novelty: 9.0 | Quality: 8.0

Read Paper | PDF


5. Semantic Overlays: Mitigating Prompt Injection with Annotations Beyond Tokens and Steering Vectors

Joshua Penman

Why it matters: Semantic Overlays offer a promising defense against prompt injection by giving models an out-of-band representation of span identity and reporting substantial attack-reduction results.

Introduces Semantic Overlays: learned residual-stream annotations that label untrusted spans as non-executable, mitigating prompt injection while preserving readability and utility across multiple benchmarks.

Score: 8.7/10 | Significance: 9.0 | Novelty: 9.0 | Quality: 8.0

Read Paper | PDF


Honorable Mentions

  • Beyond Vector Hiding: Breaking and Mitigating Shared-Direction Weight Obfuscation in TEE-Offloaded Large Language Models - The work exposes practical model-recovery failures in a proposed TEE offloading defense and provides a stronger masking design for securing LLM deployment.
  • What Guides the Agent? Adjudicating Unauthorized Behavior via Localizing Behavior-Guiding Instructions - A runtime attention-based guardrail identifies which instructions drive agent tool calls, helping defend against indirect prompt injection and tool poisoning.
  • When Context Gets Root: Privilege Escalation in LLM Harnesses - Shows that agent harnesses can defeat instruction hierarchies by elevating attacker-controlled content, exposing a broad and consequential control failure in coding agents.
  • The Framing Gap: Indirect Prompt-Injection Exfiltration Defeats Surface-Level Defenses in Tool-Using Agents - It demonstrates that tool-agent security depends more on architectural capability isolation than on asking the model to recognize cleverly reframed attacks.
  • ROPE: Routed Origin Policy Enforcement against Indirect Prompt Injection - ROPE offers a principled, high-utility defense against indirect prompt injection in increasingly autonomous tool-using agents.

This digest reviewed 1390 papers and selected the top 10 for their significance to AI safety research.

View all papers on The Guardrail


The Guardrail: Curated AI Safety Research from arXiv

Don't miss what's next. Subscribe to The Guardrail - AI Safety Weekly Research Digest:
← Newer The Guardrail Weekly Digest: 2026-08-31 - 2026-09-06 Older → The Guardrail Weekly Digest: 2026-08-17 - 2026-08-23
Twitter
LinkedIn
Powered by Buttondown, the easiest way to start and grow your newsletter.