The Commonplace: Generative AI in Action: Field Experimental Evidence from…
|
The Commonplace
Weekly Research Digest · August 04, 2026
|
The Delta
Coming in, Governance & Regulation leaned positive (661 papers); this week, a counter-signal appears.
What Moved & What Held
Coming in, the standing view was: GenAI reliably boosts speed and standardizes work in structured, repeatable tasks, disproportionately helping lower‑performing workers; decision quality and persuasion effects are mixed and design‑sensitive; audit and governance tools often miss joint behavior; and infra advances keep shifting deployment costs.
This week adds two strong field experiments that move the productivity story toward established in the studied service and hiring funnels, plus cleaner estimates that AI tone and credibility shape downstream choices in ways not captured by capability benchmarks. On governance, formal results clarify why single‑agent price‑level audits miss communication‑free collusion that preserves marginals. Still holds this week: benefits concentrate in routine workflows, quality effects vary by task and user, learning can erode without engagement, and audit/provenance design remains a first‑order risk.
Top Papers
Key: each paper is tagged Relation (New, Confirms, Extends, Tension, Challenges) and evidence status (established, suggestive, framework, descriptive). Study design (RCT, quasi-experiment) is shown separately in parentheses. full key
Confirmsestablished
Generative AI in Action: Field Experimental Evidence from Alibaba's Customer Service Operations
Xiao Ni, Yiwei Wang, Tianjun Feng, Lauren Xiaoyan Lu, Yitong Wang, Congyi Zhou; Alibaba collaboration; randomized field experiment
In Alibaba’s China-based customer support operation, random access to a GenAI assistant speeds chats and raises customer ratings, with the largest gains among low‑baseline agents; objective resolution quality does not fall. This randomized controlled trial (RCT) in live production aligns with prior lab-in-the-field findings that GenAI reduces variance and lifts the lower tail.
So what: If this holds, the risk you own is overindexing on averages and missing that returns come from variance reduction, not across‑the‑board uplift.
Confirmsestablished
Voice AI in Firms: A Natural Field Experiment on Automated Job Interviews
Brian Jabarian, Luca Henkel; natural field experiment with randomized interviewer assignment
Among 70,884 applications, assignment to automated voice interviews increases offers by about 12% and starts or early retention by about 18%, with no detectable productivity drop among hires who came through AI‑led screens; human reviewers still made final decisions. This strengthens the case that structured, consistent elicitation improves selection without short‑run output penalties in the studied funnels.
So what: If this holds, the risk you own is misattributing better funnel yield to looser standards when the gain is coming from process consistency.
Extendsestablished
John Conlon, Peter Schwardmann; preregistered randomized experiment
In an RCT with 1,500 participants across 30 incentivized decision tasks, a sycophantic LLM still depolarizes choices on average (about 0.22 standard deviations toward center) relative to no‑chat controls, with higher sycophancy attenuating that depolarization. This nuances the standing view that agreement‑seeking AI primarily amplifies priors.
So what: If this generalizes, the risk you own is assuming any one alignment tweak (less sycophancy) monotonically improves decision quality when effects depend on task and baseline tilt.
Also Notable
Newframework Collusion with Competitive Marginals: Price-Level Audits Are Blind by Construction (Xin Xu, Chengrui Wu, Jiayu Lu, Kaizhen Tan, Siru Tao, Hanzhe Hong): Formally derives and illustrates that conspiracies preserving each bidder’s marginal price distribution can evade single‑agent audits, indicating single‑agent tests may miss such coordination without joint‑dependence checks.
Confirmsestablished Faster, Higher, Stronger? The Impact of GenAI on Knowledge Work Productivity - Evidence from the Field (Bottesch, Sven; Schwenke, Chiara; Zimmermann, Jakob; Förster, Maximilian; Klier, Mathias): Randomized controlled trial (N=128) finds GenAI speeds tasks broadly, improves packaging and creation quality, but reduces acquisition‑task quality, again with larger lifts for low performers.
Newdescriptive The Tokenizer Tax: Quantifying and Explaining the Cross-Lingual Cost of Subword Tokenization for Indian Languages (Priyansh Srivastava): Measures an average 8x token cost versus English for common tokenizers, shrinking usable context in Indian languages absent multilingual tokenizers.
Newdescriptive OptiRepair: Closed-Loop Diagnosis and Repair of Supply Chain Optimization Models with LLM Agents (Ruicheng Ao, David Simchi-Levi, Xinshang Wang): Purpose‑trained 8‑billion‑parameter (8B) models repair infeasible supply‑chain linear programs (LPs) far more often than application programming interface (API)‑based LLMs in tests, suggesting targeted small models could be more effective for operational tooling in this setting.
Extendsestablished Perceived Political Bias in LLMs Reduces Persuasive Abilities (Matthew DiGiuseppe, Joshua Robison): US preregistered survey experiment (N=2,144) finds bias warnings cut LLM persuasion by about 28%, quantifying credibility’s role in uptake.
Newframework Winner's Curse Drives False Promises in Data-Driven Decisions: A Case Study in Refugee Matching (Hamsa Bastani, Osbert Bastani, Bryce McLaughlin): Theory and calibrations suggest model‑based policy evaluation can overstate gains even when causal benefit is zero because of a winner’s curse.
Newframework The Condensate Theorem: Transformers are O(n), Not O(n^2) (Jorge L. Ruiz Williams): Presents a theorem and tests suggesting attention might be projected onto a learned manifold with bit‑exact outputs in demonstrations, enabling long‑context speedups in those tests.
Confirmsestablished The Persuasion Paradox: When LLM Explanations Fail to Improve Human-AI Team Performance (Ruth Cohen, Lu Feng, Ayala Bloch, Sarit Kraus): Controlled studies find fluent explanations raise confidence and reliance without improving accuracy, and can hinder error recovery.
Newdescriptive Training LLMs with Fault Tolerant HSDP on 100,000 GPUs (Omkar Salpekar, Rohan Varma, Kenny Yu, Vladimir Ivanov, Yang Wang, Ahmed Sharif, Min Si, Shawn Xu, Feng Tian, Shengbao Zheng, Tristan Rice, Ankush Garg, Shangfu Peng, Shreyas Siravara, Wenyin Fu, Rodrigo de Castro, Adithya Gangidi, Andrey Obraztsov, Sharan Narang, Sergey Edunov, Maxim Naumov, Chunqiang Tang, Mathew Oldham): System design reports about half the failure stalls and roughly double effective utilization at roughly 100,000 graphics processing units (GPUs) in their tests, with no measured accuracy loss.
Confirmsestablished How AI Impacts Skill Formation (Judy Hanwen Shen, Alex Tamkin): RCTs find AI coding assistance can reduce novices’ conceptual learning and debugging skill even when productivity rises.
Newsuggestive Bayesian and Motivated Reasoning in AI Agents (Eddie Yang): Agent outputs shift with framing consistent with prior‑aligned motivated reasoning, raising reliability questions in consequential domains.
Newdescriptive Artificial Intelligence: Supply-Chain Chokepoints and the Reach of Industrial Policy (Piyush Akimitsu): Maps steep concentration upstream (packaging, lithography, mineral refining) relative to models and cloud using the Herfindahl‑Hirschman Index (HHI).
Extendssuggestive A robust association between LLM use and scientific productivity: Assessing stopping-time selection (Keigo Kusumegi, Xinyu Yang, Paul Ginsparg, Mathijs de Vaan, Toby Stuart, Yian Yin): After calibrating stopping‑time bias, detected LLM use on arXiv remains associated with higher author output.
Newdescriptive FairFund-Bench: Evaluating Distributive Bias in LLM Resource Allocation (Martin Lukk): Allocation bias estimates flip with audit format; causal-need framing appears larger than demographic effects across tested protocols.
Newdescriptive Can LLM Agents Price Competitively? A Dynamic Multi-Attribute Auction Benchmark for Agentic Commerce (Shimaa Ahmed, Yiwei Cai, Mohsen Minaei, Rahul Rachuri): Agents win customers but underperform on profit; margin‑per‑win and post‑shock adaptation matter more than simple win rates in this benchmark.
Newdescriptive Hidden Errors in Big Data: The Case of Property Records (Evelyn Smith, Emma Harvey, Jacob Goldin, Daniel E. Ho): Brokered property datasets show 1–2% large price errors and 12–15% coverage gaps versus county records, shifting tax and regressivity estimates in this sample.
Newdescriptive Share the Judge, Learn the Deferral: Where Specialization Helps LLM Evaluation (Weining Zhang): Shared‑judge initialization plus rubric routing preserves evaluator accuracy and coverage better than isolated small specialists.
Challengesdescriptive When Synthetic Users Fail: A Cross-Domain Benchmark of LLM-Simulated Human Survey Responses (Zihan Chen, Di Zhu, Lei Nico Zheng): LLMs underperform demographic baselines at predicting individual survey responses and exaggerate between‑group differences under tested setups.
Newdescriptive A Density-Matrix Framework for Electronic-Structure Analysis of Functional-Group and Salt Effects in Lithium-Metal Electrolytes (Mingkang Liu, Huize Yu, Yanbin Gao, Nan Yao, Xiang Chen, Lei Shen): ML‑accelerated density‑matrix predictions scale electronic‑structure screening to hundreds of thousands of molecules.
Newdescriptive HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following (Liudas Panavas, Sebastian Minus, Bradley Monton, Derek Ray, Suhaas Garre, Sushant Mehta, Edwin Chen): Agents show low strict pass rates on long policy‑following tasks with near‑miss and false‑compliance failure modes.
Newdescriptive Instruction-Tuned Language Models Cannot Sample from Distributions They Can Describe (Chaemin Jang, Dongman Lee, Jihee Kim): Instruction tuning often induces output collapse instead of proper sampling, undermining synthetic‑sampling use cases.
Newestablished Integrating Predictive Models into Two-Sided Recommendations: A Matching-Theoretic Approach (Kazuki Sekiya, Suguru Otani, Yuki Komatsu, Sachio Ohkawa, Shunya Noda): Theory, calibrated simulations, and a field trial indicate exposure‑constrained deferred acceptance can improve effective matches by reducing congestion.
Confirmsestablished AI-Mediated Feedback Improves Student Revisions: A Randomized Trial with FeedbackWriter in a Large Undergraduate Course (Xinyi Lu, Kexin Phyllis Ju, Mitchell Dudley, Larissa Sano, Xu Wang): Providing teaching assistants (TAs) with LLM suggestions improves essay revisions; effects scale with TA adoption.
Newsuggestive Thin Keys, Full Values: Reducing KV Cache via Low-Dimensional Attention Selection (Hengshuai Yao, Xing Chen, Ahmed Murtadha, Guan Wang): Factored keys compress the key‑value (KV) cache substantially with modest pretraining cost and small quality loss in 7‑billion‑parameter (7B) models.
Extendsestablished When Algorithms Meet Ethics: Systematic Evidence of Framing Effects in LLM Organizational Decision-Making (Jonathan H. Westover): Preregistered factorial RCT (14,306 responses) finds large framing effects (scarcity, severity, procedural justice) on LLM ethical recommendations.
Extendssuggestive Choose Your Agent: Tradeoffs in Adopting AI Advisors, Coaches, and Delegates in Multi-Party Negotiation (Kehang Zhu, Nithum Thain, Vivian Tsai, James Wexler, Crystal Qian): Delegation increases surplus most in their setup, but users prefer advisory control, dampening realized gains.
Extendsestablished RELATE: A Reinforcement Learning-Enhanced LLM Framework for Advertising Text Generation (Jinfang Wang, Jiajie Liu, Jianwei Wu, Ziqin Luo, Zhen Chen, Chunlei Li, Biao Han, Tao Deng, Yi Li, Shuanglong Li, Lin Liu): Joint reinforcement learning (RL) for conversion and compliance improves ad performance in offline tests and online A/B tests.
What Moved
Contested & Watch
Methods Spotlight
Fault‑tolerant hybrid shared data parallelism (FT‑HSDP): Training LLMs with Fault Tolerant HSDP on 100,000 GPUs reports roughly double effective utilization at extreme scale in testing, with no measured accuracy loss, shifting the economics of large‑model training.
Condensate manifold projection for attention: The Condensate Theorem proposes a provable approach, with demonstrations suggesting bit‑exact long‑context speedups by projecting attention onto a learned topology.
Closed‑loop LP repair with solver‑verified rewards: OptiRepair couples infeasibility diagnosis with targeted training of an 8B model, materially outperforming generic APIs on repairing supply‑chain optimization models in evaluated cases.