The Guardrail - AI Safety Weekly Research Digest logo

The Guardrail - AI Safety Weekly Research Digest

Archives
Log in
Subscribe
17 August 2026

The Guardrail Weekly Digest: 2026-08-10 - 2026-08-16

Weekly Digest

The Guardrail Weekly Digest

Week of 2026-08-10 to 2026-08-16

10 papers were selected from 1,130 reviewed this week, spanning governance, alignment, and robustness. Deployment Decision Reliability proposes a generalizability-theory framework for sizing long-horizon agent evaluations, addressing the risk of overtrusting unreliable leaderboards in deployment decisions. Several papers identify severe model-security failures: Once Poisoned, Arbitrarily Controlled demonstrates programmable VLM backdoors, Stealing Reasoning Traces from Proprietary LLM APIs shows how encrypted traces can enable extraction and leakage, and Diffusion LLMs as Targets and Adversaries exposes transferable safety bypasses. Genotypic Triggers extends these concerns to drug discovery by revealing host-specific backdoors that can conceal pharmacogenomic hazards in antimicrobial peptide models.


Top Papers This Week

1. Deployment Decision Reliability: A Generalizability-Theory Framework for Sizing Long-Horizon Agent Evaluations

Vasundra Srinivasan

Why it matters: DDR helps enterprises avoid overtrusting unreliable agent leaderboards when making high-stakes deployment decisions.

A Generalizability-Theory analysis finds agent leaderboards largely reflect agent-by-task specialization rather than general capability, with reliability collapsing on hard tasks. DDR converts variance estimates into deployment decisions and reporting guidance.

Score: 8.7/10 | Significance: 9.0 | Novelty: 8.0 | Quality: 9.0

Read Paper | PDF


2. Once Poisoned, Arbitrarily Controlled: A Programmable Backdoor in VLMs

Tao Lin, Gaojie Jin, Zongxin Liu...

Why it matters: This work exposes a severe VLM supply-chain and deployment risk: one poisoning stage can enable flexible, stealthy control over arbitrary future outputs.

Introduces a programmable VLM backdoor that maps attacker-chosen captions to stealthy image triggers after poisoning, enabling unseen target control without retraining. Experiments show high attack success, preserved clean utility, and resistance to classical defenses.

Score: 8.7/10 | Significance: 9.0 | Novelty: 9.0 | Quality: 8.0

Read Paper | PDF


3. Stealing Reasoning Traces from Proprietary LLM APIs

Alexander Panfilov, David Schmotz, Ilia Shumailov...

Why it matters: A serious architectural flaw turns encrypted reasoning traces into a channel for model extraction, sensitive-data leakage, hazardous-content disclosure, and stealthy agent attacks.

Exposes a cross-model decryption jailbreak that extracts proprietary reasoning, PII, credentials, and hazardous content from encrypted client-side traces, while enabling invisible prompt injection into agentic rollouts; proposes cryptographic and system-level mitigations.

Score: 8.75/10 | Significance: 9.0 | Novelty: 9.0 | Quality: 8.0

Read Paper | PDF


4. Diffusion LLMs as Targets and Adversaries: Mechanistic Safety Exploits

Elena Dumitrescu, Gert Lek, Lydia Y. Chen...

Why it matters: The work reveals that diffusion LLM safety mechanisms can transfer across architectures and be efficiently bypassed, highlighting a major security gap in emerging model designs.

Exposes sparse, transferable safety mechanisms in diffusion LLMs and introduces SN-Guided Diffusion, an offline jailbreak that prunes or steers safety neurons, achieving high transfer attack success at low generation cost.

Score: 8.7/10 | Significance: 9.0 | Novelty: 9.0 | Quality: 8.0

Read Paper | PDF


5. Genotypic Triggers: Exposing Pharmacogenomic Blind Spots via Host-Specific Backdoors in Generative Antimicrobial Peptide Models

Doniyorkhon Obidov, Xiaolong Guo, Yonghui Li...

Why it matters: Shows how peptide-generation models can conceal genotype-specific health hazards, exposing a serious blind spot in AI-enabled drug-discovery safety pipelines.

Introduces Genotypic Triggers, backdoors that make peptide generators produce immunogenic antimicrobial peptides for carriers of targeted HLA alleles while preserving potency and passing conventional safety screens.

Score: 8.7/10 | Significance: 9.0 | Novelty: 9.0 | Quality: 8.0

Read Paper | PDF


Honorable Mentions

  • From Prompt Injection to Web Exploitation: Revisiting Classic Vulnerabilities in LLM-Integrated Applications - This work reframes prompt injection as a bridge from attacker input to classic web exploits, giving developers a concrete taxonomy and mitigation framework for LLM-integrated systems.
  • "Operator, can you hear me?" A Faithful Line into the UNISOC Baseband - A faithful baseband re-host makes previously inaccessible cellular firmware security analysis practical, with implications for modem and automotive-system security.
  • IO Factory: Simulating AI-Enabled Influence Campaigns at Scale - A scalable, inspectable testbed for studying and red-teaming coordinated AI influence operations before they occur in real platforms.
  • SteerBench-Work: A Benchmark for Agent Steering at Action Boundaries - A consequential-action benchmark reveals that capable workplace agents often over-refuse—and can miss reversed safety evidence—making steering calibration a distinct deployment risk.
  • Jagged Judges: Epistemic Stability Under Silence, Pressure, and Persistence - Reliable LLM judges are foundational to evaluation and reward modeling, yet this work shows they can be systematically destabilized and corrupted by pressure.

This digest reviewed 1130 papers and selected the top 10 for their significance to AI safety research.

View all papers on The Guardrail


The Guardrail: Curated AI Safety Research from arXiv

Don't miss what's next. Subscribe to The Guardrail - AI Safety Weekly Research Digest:
Older → The Guardrail Weekly Digest: 2026-08-03 - 2026-08-09
Twitter
LinkedIn
Powered by Buttondown, the easiest way to start and grow your newsletter.