The Guardrail - AI Safety Weekly Research Digest logo

The Guardrail - AI Safety Weekly Research Digest

Archives
Log in
Subscribe
22 June 2026

The Guardrail Weekly Digest: 2026-06-15 - 2026-06-21

Weekly Digest

The Guardrail Weekly Digest

Week of 2026-06-15 to 2026-06-21

This week we reviewed 1259 papers and selected the top 10 for their significance to AI safety research.


Top Papers This Week

1. Rift: A Conflict Signature for Deception in Language Models

Petr Nyoma

Why it matters: Rift identifies a robust, cross-model internal 'conflict signature' that allows for the near-perfect detection of deceptive behavior in LLMs, providing a critical tool for mechanistic interpretability.

Rift identifies a "conflict signature"—elevated residual rank—that distinguishes deceptive model outputs from honest errors or hallucinations. This provides a robust, cross-architecture, and zero-shot method to detect latent deception, addressing a critical gap in ELK.

Score: 9.0/10 | Significance: 10.0 | Novelty: 10.0 | Quality: 9.0

Read Paper | PDF


2. AI systems out-persuade expert humans

Kobi Hackenburg, Caroline Wagner, Luke Hewitt...

Why it matters: This study provides empirical evidence that frontier AI models can consistently outperform expert human persuaders, highlighting urgent risks regarding AI-driven manipulation and democratic discourse.

Frontier AI models consistently outperform expert human debaters in persuasion, even when humans are incentivized and coached. This highlights a critical safety risk: AI’s ability to rapidly deploy information can be weaponized to manipulate public opinion at scale.

Score: 10.0/10 | Significance: 10.0 | Novelty: 9.0 | Quality: 9.0

Read Paper | PDF


3. Security Engineering of OpenClaw: Analyzing Attack Surface Expansion and Trust-Boundary Violations

Saeid Jamshidi, Arghavan Moradi Dakhel, Kawser Wazed Nafi...

Why it matters: This paper provides a rigorous, quantitative framework for measuring how multi-agent system architecture—independent of model intelligence—exponentially expands attack surfaces and privilege drift.

OpenClaw demonstrates that multi-agent systems exhibit non-linear security degradation, with compromise probability jumping from 0.24 to 0.86 as agent count increases. This highlights that structural aggregation, not just model capability, drives privilege drift and risk.

Score: 9.0/10 | Significance: 10.0 | Novelty: 9.0 | Quality: 9.0

Read Paper | PDF


4. Vision-language models for chest radiography do not always need the image

Mahshad Lotfinia, Sebastian Ziegelmayer, Lisa Adams...

Why it matters: This paper exposes a critical failure mode in medical AI where models achieve high accuracy by exploiting linguistic priors rather than visual evidence, demonstrating that accuracy metrics are insufficient for clinical safety.

A causal audit of medical VLMs reveals that many models rely on text-based finding-name priors rather than visual evidence. By demonstrating that accuracy metrics fail to distinguish grounded reasoning from shortcut learning, the study mandates grounding audits for safety.

Score: 9.0/10 | Significance: 9.0 | Novelty: 9.0 | Quality: 9.0

Read Paper | PDF


5. MIRAGE: Auditing Anti-Muslim Bias in Frontier LLMs Across Reasoning, Agentic, and Time-Coupled Conditions

Noor Islam S. Mohammad, Tamim Sheikh

Why it matters: MIRAGE exposes how frontier LLMs exhibit amplified anti-Muslim bias in complex, agentic, and reasoning-heavy workflows that standard static benchmarks fail to capture.

MIRAGE reveals that CoT reasoning and agentic workflows amplify anti-Muslim bias in frontier LLMs, with biases persisting despite standard prompt-based mitigations. This highlights a critical failure in current safety evaluations that ignore complex, multi-step deployment.

Score: 9.0/10 | Significance: 9.0 | Novelty: 9.0 | Quality: 9.0

Read Paper | PDF


Honorable Mentions

  • Incentives and Evidence in Learned Service Orchestration - A rigorous meta-analysis revealing that the field of learned service orchestration is built on fragile benchmarks and misaligned incentives, providing a blueprint for more robust AI evaluation practices.
  • The Gate Is Only as Honest as Its Contracts: ContractGuard for the Contract Layer of Risk-Aware Causal Gating - ContractGuard addresses a critical vulnerability in tool-augmented agents by securing the integrity of tool contracts, shifting the defense from fragile prompt-based filtering to robust, verifiable provenance.
  • DeFAb: A Verifiable Benchmark for Defeasible Abduction in Foundation Models - DeFAb introduces a rigorous, verifiable benchmark for defeasible abduction that exposes the fundamental inability of frontier models to perform consistent logical reasoning, providing a clear path for formal-verification-based alignment.
  • SafeClawBench: Separating Semantic, Audit-Evidence, and Sandbox Harm in Tool-Using LLM Agents - SafeClawBench provides a critical, multi-stage evaluation framework that decouples semantic compliance from actual executable harm in tool-using AI agents.
  • Towards Understanding and Measuring COGNITIVE ATROPHY in LLM Behaviour - This paper introduces 'Cognitive Atrophy' as a critical new metric for evaluating how AI-mediated mental health support may inadvertently foster user dependency rather than autonomy.

This digest reviewed 1259 papers and selected the top 10 for their significance to AI safety research.

View all papers on The Guardrail


The Guardrail: Curated AI Safety Research from arXiv

Don't miss what's next. Subscribe to The Guardrail - AI Safety Weekly Research Digest:
← Newer The Guardrail Weekly Digest: 2026-06-22 - 2026-06-28 Older → The Guardrail Weekly Digest: 2026-06-08 - 2026-06-14
Twitter
LinkedIn
Powered by Buttondown, the easiest way to start and grow your newsletter.