The Commonplace: Whose doctor does the AI recommend? An algorithm audit of…
|
The Commonplace
Weekly Research Digest · August 18, 2026
|
The Delta
Coming in, Governance & Regulation leaned positive (661 papers); this week, a counter-signal appears.
What Moved & What Held
Coming in, the standing view was that LLMs and agents change choices and workflows, with partial behavioral pass-through and measurable firm-side gains in field A/Bs; governance can reduce visible harms but adds friction; and one-shot audits and thin disclosures often misstate safety, fairness, or procurement risk.
This week adds causal magnitudes in two high-stakes settings and sharpens trade-offs: a lab-in-the-field pension experiment estimates pass-through at 37% without Sharpe improvements (return per unit of risk), and an LLM doctor-recommender audit estimates dollar-equivalent sizes on ordering and demographic effects; on the production side, a causal targeting system increases incremental value while a guardrail rollout reports large reductions in hallucinations at the cost of slower tasks. Still holds this week: AI steers decisions and can lift firm outcomes, but governance that relies on single snapshots or unspecified metrics can be unreliable.
Top Papers
Key: each paper is tagged Relation (New, Confirms, Extends, Tension, Challenges) and evidence status (established, suggestive, framework, descriptive). Study design (RCT, quasi-experiment) is shown separately in parentheses. full key
Newestablished
Syeda Anshrah Gillani, Mirza Samad Ahmed Baig
pre-registered randomized conjoint, a choice experiment with randomized attributes
Across seven commercial LLMs and 3,024 randomized choice sets (40,068 model responses), higher ratings and lower fees increase recommendation probability, while randomized ordering and name cues induce smaller but consistent shifts worth roughly an $11 fee change. This identifies causal attribute effects on LLM-mediated recommendations in a synthetic US-style physician-card setting, filling in magnitudes the baseline lacked.
So what: If this generalizes, LLM-mediated referrals may amplify reputational inequality and embed subtle ordering and demographic skews that change who gets work.
Extendsestablished
Do People Follow AI Advice? Evidence from a Pension Portfolio Choice Experiment
Hongseok Choi, Jeongbin Kim, Matthew Kovach, Kyu-Min Lee, Euncheol Shin, Hector Tzavellas
incentivized RCT
In an RCT with N=400 pension participants, randomized aggressive vs conservative AI recommendations shift about 37% of the allocation gap into final portfolios, increasing expected returns and risk but leaving Sharpe ratios unchanged and reducing diversification. This extends prior pass-through evidence to retirement finance with causal estimates on risk exposure.
So what: If this holds, AI advice can move household risk loads without improving risk-adjusted performance, exposing providers and plans to distributional risk they may be undercounting.
Extendsestablished
From prediction to incrementality: Causal optimization for large-scale targeting and recommendation
Changshuai Wei, John Bencina, Phuc Nguyen, Andre Assuncao Silva T Ribeiro, Benjamin Zelditch
LinkedIn, production A/B test
A production system at US-based LinkedIn that combines causal lift estimation, exploration, and constrained allocation increases incremental marketing long-term value by 7.20% in an A/B test, outperforming prediction-only baselines. This adds to firm-side evidence that optimizing for incremental impact can deliver gains beyond engagement proxies.
So what: If this holds, firms relying on predictive scores risk wasting spend on users who would convert anyway and understate true return on model-driven outreach.
Also Notable
Newsuggestive Governing generative AI in organizations: a design theory and quasi-experimental field study of sociotechnical guardrails Maikel Leon - A stepped-wedge rollout (phased rollout across units over time) at a Fortune 500 firm associates layered guardrails with large drops in hallucinations and stronger audit trails, alongside a measurable time penalty and some circumvention.
Newestablished Self-evolving agentic customer support system at LinkedIn Chih Hui Wang, Mengdie Tu, Qianyun Zhang, Wei Wu, Lili Zhou, Mingqi Shen, Changshuai Wei (LinkedIn) - A two-week production A/B finds modular agentic workflows increase self-serve and measurably improve routing accuracy in this platform's sample.
Newdescriptive No task fails every time: Why one-shot audits are structurally blind to agent damage Shiven Khurdi - Repeated state-diff audits (comparing system or environment state before and after runs) find stochastic, irreversible damage events across models, implying single-run checks will often miss dangerous behaviors.
Newestablished Toward meaningful transparency for AI chatbots: Disclosing persuasive intent reduces persuasion Adrian Rauchfleisch, Andreas Jungherr - A 1,500-person RCT finds intent disclosure halves a chatbot's persuasive effect, while generic AI identity labels do not.
Newdescriptive Same system, opposite verdicts: Metric discretion in AI ethics audits and the limits of disclosure Shay Tsaban - A multiverse audit (analyzing many plausible metric and sample choices) documents how plausible metric and population choices flip fairness verdicts and are rarely disclosed.
Newdescriptive VAKRA: Evaluating multi-hop reasoning across APIs and retrieval under tool-use policies Ankita Rajaram Naik, Anupama Murthi, Benjamin Elder, Siyu Huo, Raavi Gupta, Abhinav Jain, Praveen Venkateswaran, Abdulhamid Adebayo, Danish Contractor - Accuracy drops with compositional depth and policy constraints, flagging limits for enterprise agents.
Newframework Deployment decision reliability: A generalizability-theory framework for sizing long-horizon agent evaluations Vasundra Srinivasan - Variance decompositions suggest agent-by-task interactions may dominate, so leaderboards may reflect specialization rather than general capability in this framework's samples.
Newdescriptive Frontier AI forecasting has a measurement problem: An audit of progress evidence Fabricio F Costa - Public records on compute and capabilities are sparse, concentrated, and rarely linked, which weakens naive extrapolation.
Newdescriptive Pricing the risk of runtime compression: Anytime-valid admission and a served-output law for compressed serving state Fanzhe Wei, Li Liu - In this deployment, an admission ledger makes compression risk auditable and fallbacks run at about half at matched loss on a live mixture-of-experts stack (a model with multiple specialized subnetworks).
Newdescriptive Human versus computer vision Elena Sirotkina - On news photos, a simple center bias beats trained saliency models, and residuals vary with viewer demographics.
What Moved
Contested & Watch
Methods Spotlight
Pre-registered randomized conjoint audit across multiple LLMs: Gillani and Baig. Causally identifies how ratings, fees, names, and ordering drive LLM recommendations at scale, yielding transparent marginal effects with clustered uncertainty.
End-to-end causal uplift optimization with exploration and constrained allocation: Wei et al. (LinkedIn). Puts modern causal inference into production, showing incremental value gains over prediction-only systems in a live A/B.
Repeated ground-truth state-diff auditing for agents: Khurdi. Captures stochastic, irreversible harms that one-shot audits miss, a practical template for safety-critical evaluations.