The Commonplace: AI-assisted teams outperform AI-led teams but not…
|
The Commonplace
Weekly Research Digest · July 13, 2026
|
From Alex
Correction: In last week's author note, I described the linked NFW Reader piece incorrectly. It was about using Copilot in Word for causal-impact work, not Copilot Studio. The link itself was correct. I also included a reader-facing link for the OES Dashboard that did not resolve publicly; the correct link is the OES Dashboard. The archive has been corrected.
Most discussion of AI agents starts with worker productivity. "The Agentic Economy" asks a second economic question: what changes when assistant agents and service agents can communicate directly? The visual Reader piece maps the lower communication costs that could follow, the choice between platform-run walled gardens and an open web of agents, and six markets that could shift.
The Delta
Coming in, Research Productivity leaned positive (373 papers); this week, a counter-signal appears.
What Moved & What Held
Coming in, the standing view was that AI works best as an augmentor alongside skilled humans, not a drop-in replacement; early labor-market impacts look limited and heterogeneous; firm-level productivity gains appear in pockets while macro productivity remains uncertain; and AI intermediaries may redirect value in digital markets.
This week adds hard causal evidence that AI-led workflows can fail on critical error detection even when AI assistance helps, national administrative data that add weight against near-term displacement fears for early-career workers, and estimated magnitudes for how AI search compresses outbound web referrals. Orchestration and memory studies report sizable cost and latency improvements without observed quality loss in tested tasks, and in South Korea, hours adjustments appear in high-exposure industries rather than headcount losses. Still holds this week: augmentation over automation, limited immediate displacement, uneven firm gains, and unresolved micro-to-macro scaling.
Top Papers
Key: each paper is tagged Relation (New, Confirms, Extends, Tension, Challenges) and evidence status (established, suggestive, framework, descriptive). Study design (RCT, quasi-experiment) is shown separately in parentheses. full key
Confirmsestablished
Abel Brodeur, David Valenta, Alexandru Marcoci, Juan P. Aparicio, Derek Mikola, Bruno Barbarioli, Rohan Alexander, Lachlan Deer, Tom Stafford, Lars Vilhuber, Gunther Bensch, Fabio Motoki, Mohamed Abdelhady, Yousra Abdelmoula, Ghina Abdul Baki, Tomás Aguirre, Sriraj Aiyer, Shumi Akhtar, Farida Akhtar, Melle R. Albada, Micah Altman, David Angenendt, Zahra Arjmandi Lari, Jorge Armando De Leon Tejada, David Rodriguez Arana, Igor Asanov, Anastasiya-Mariya Noha, Rebecca Ashong, Tobias Auer, Francisco J. Bahamonde-Birke, Bradley J. Baker, Söhnke M. Bartram, Dongqi Bao, Lucija Batinovic, Tommaso Batistoni, Monica Beeder, Louis‐Philippe Beland, Carsten Gero Bienz, Christ Billy Aryanto, Cylcia Bolibaugh, Carl Bonander, Ramiro Bravo, Egor Bronnikov, Stephan Bruns, Nino Buliskeria, Sara Caicedo-Silva, Andrea Calef, Juan Sebastian Cano Arias, Gustavo A. Castillo Alvarez, Solomon Caulker, Simonas Cepenas, Arthur Chatton, Zirou Chen, Ngozi Chioma Ewurum, Anda-Bianca Ciocîrlan, Felix J. Clouth, Jason Collins, Nikolai Cook, Cesar Cornejo, Joao Craveiro, Jonathan Crechet, Jing Cui, Niveditha Chalil Vayalabron, Christian Czymara, Carlos Daniel Bermúdez Jaramillo, Hannes Datta, Lien Denoo, Arshia Dhaliwal, Nency Dhameja, Elodie Djemai, Erwan Dujeancourt, Uğurcan Dündar, Thibaut Duprey, Yasmine Eissa, Youssef El Fassi, Ismail El Fassi, Keaton Ellis, Ali Elminejad, Mahmoud Elsherif, Aysil Emirmahmutoglu, Giulian Etingin-Frati, Emeka Eze, Jan Fabian Dollbaum, Jan Feld, Andres Felipe Rengifo Jaramillo, Guidon Fenig, Victoria Fernandes, Lenka Fiala, Lukas Fink, Mojtaba Firouzjaeiangalougah, Sara Fish, Jack Fitzgerald, Rachel Forshaw, Alexandre Fortier-Chouinard, Louis Fréget, Joris Frese, Jacopo Gabani, Sebastián Gallegos, Max C. Gamill, Attila Gáspár; randomized controlled trial
In a randomized controlled trial that assigned 288 researchers into 103 teams across human-only, AI-assisted, and AI-led arms, AI assistance matched human reviewers on overall reproducibility while AI-led teams reproduced only about 37% and missed more major coding errors. This supports the standing view that augmentation helps but autonomy degrades quality on complex verification tasks in this task.
So what: If this holds, delegating high-stakes verification to AI carries higher undetected error risk and reputational exposure than many current quality-assurance models assume.
Confirmssuggestive
Labor Market Consequences of Generative AI: Early Evidence from Norway
Dennis Facius, R. Iacono; quasi-experiment using population registers
Using Norwegian population registers from 2015 to March 2025 and multiple quasi-experimental strategies (within-firm composition difference-in-differences, occupation-level synthetic difference-in-differences, and firm-level shift-share) that exploit the November 2022 ChatGPT release, the authors find no robust displacement of early-career workers in highly exposed occupations and no clear effects across other cohorts. This aligns with the baseline that near-term employment effects are muted and heterogeneous.
So what: If this holds, overestimating immediate displacement risk could misdirect training budgets and social insurance planning toward the wrong time horizon.
Confirmssuggestive
Answering Without Referring: How AI Search Rewrites the Web's Economic Bargain
Qiaoni Shi, Kai Zhu, Kai Gu; quasi-experimental clickstream study
US Comscore desktop clickstream data indicate ChatGPT produces outbound clicks in only 5.2% of conversation sessions and that broader access to ChatGPT Search is associated with about 9.4% lower traditional search use during rollout, with clicks skewing to specialized sites rather than ad-heavy portals. This provides quantitative backing for the claim that AI intermediaries compress referral flows that fund the open web.
So what: If this generalizes, referral-dependent publishers and ad markets face sharper traffic and revenue compression than current operating plans contemplate.
Also Notable
Confirmsestablished Experimental Evidence on the Learning Impact of Generative AI (Zara Contractor, Germán Reyes), Undergraduate randomized controlled trial finds AI access lifts immediate test scores by 0.27 SD with one-week persistence concentrated among students who use AI for explanation, consistent with augmentation over automation in learning.
Extendssuggestive AI, Output, and Employment (Andrew Johnston, Christos A. Makridis), US administrative employer data associate higher occupational AI exposure with about 7% higher output and 4% higher employment where collaboration is essential, alongside a lower labor share, pointing to distributional shifts even when headcounts grow.
Extendssuggestive Artificial Intelligence Exposure and Working Hours: Evidence From South Korea (Taiwon Ha), Industry-level exposure measures link higher AI exposure to larger post-2022 declines in weekly hours with no pre-trend differences, suggesting hours may adjust before employment in this context.
Newdescriptive AI Adoption in S&P 500 Firms (Yang Yu, Martin Fleming, Lucy Hampton, Christophe Combemale, Neil Thompson), A 10-K based deep-integration measure rises to 11% in 2025, quadrupling since 2022, with a J-shaped profit association and no clear capex or productivity signals in filings, documenting who is moving first.
Confirmsestablished Artificial Intelligence and the Digital Economy: Impact on Employment, Productivity, and Market Structures (Dr. Snehal Mistry, Siddharth Thakkar), Systematic review of 78 studies reports firm-level productivity gains and routine task displacement often offset by complementary roles, with uneven distributional outcomes.
Extendssuggestive The carbon reduction effect of China’s national AI innovation pilot zone policy (劉南勳, Shuqing Wang, Yuanhong Peng), Staggered city-level quasi-experiment links AI pilot-zone designation to about 6% lower urban CO2, rising to 15.6% with spatial spillovers, suggesting policy-associated environmental effects.
Extendssuggestive Artificial Intelligence and Urban Green Productivity in China: The Role of Green Computing Capacity and Transmission Channels (Xiaoxiao Tian, Wei Guo, Jingyu Liao), City panels associate AI patenting with higher green productivity where green computing capacity and industry bases are strong, highlighting the role of complements.
Newdescriptive Seeing is Free, Speaking is Not: Uncovering the True Energy Bottleneck in Edge VLM Inference (Junfei Zhan, Haoxun Shen, Mingang Guo, Zixuan Huang, Tengjiao He), On-device profiling finds output token decoding dominates energy and latency for vision-language models, shifting where efficiency gains are likely to come from.
Newsuggestive The Harness Effect: How Orchestration Design Sets the Token Economics of Enterprise Agentic AI (Muayad Sayed Ali, Aliaksandra Novik, Anji Boddupally, Artem Yavorskyi, Chris Nickerson, Daniel Rica, Emily DuGranrut, Felix Leung, Garrett Prince, Grace Barnett, Heath Robinson, Hosain Al Ahmad, Jesse Resnick, Juan Carlos Farah, Jyothi Swaroop Meruga, Leonid Kuznetsov, Luke Gorham, Marie Schmoll, Michael Paciullo, Saumya Das, Sharath Sheripally, Tommy Griscom, Mykyta Osadchyi, Neha Mantri, Nick Westrum, Olivia Benowitz, Parikshith Kulkarni, Radik Chernyshov, Rakshith Vasudev, Rohith Nadimpally, Vikas Gangadevi, Waseem AlShikh), Controlled engineering experiments report 30–45% lower token costs and latency with orchestration layers while holding measured task quality constant, pointing to a potential lever for enterprise deployments.
Newsuggestive Shared Selective Persistent Memory for Agentic LLM Systems (Sanjana Pedada, Aditya Dhavala, Neelraj Patil), A memory design achieves 96% task completion while cutting tokens and time versus naive persistence in tested tasks, suggesting architecture choices can mitigate cost-quality tradeoffs.
Tensiondescriptive 3100 Opinions on Code Review in an AI World: Building Causal Theory from Practitioner Discourse (Shyam Agarwal, Courtney Miller, Christian Kästner, Bogdan Vasilescu), Repository mining suggests agent-authored pull requests are reviewed less and merged faster, but patterns flip under alternative analytic choices, underscoring governance ambiguity around AI-generated code.
Newframework Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops (Mingguang Chen, Licheng Wang, Bo Qu), Survey distinguishes verifiable, bounded self-refinement from speculative open-ended self-improvement and argues current systems remain compute and grounding constrained.
Extendssuggestive Two-player Alternate Uses Test: A Controlled Testbed for Interactive Human-AI and Human-Human Co-Creation (Babak Hemmatian, Anita Keshmirian, Yijun Lin, Shravan Ramamoorthy, Maryam Jahadakbar, Eli Khuri-Reid, Jingtong Wang, Sarah Hadjarab, Sindre Veum, Pranav Gupta, Deepak Somaya, Lav R. Varshney), In a small controlled setting, GPT-4 partners generate originality on par with human partners, conditional on motivation and prior exposure.
Newframework Analysis of labor productivity in the context of technological transformations (T. Martyn, V. Nitsenko, V. Kyrylenko, O. Tkachenko, Y. Stavska, O. Kulhanik), Proposes an "AI phase" where organizational frictions and AI debt slow macro productivity even as micro gains accrue.
Extendssuggestive Asymmetric effects of renewable energy and artificial intelligence on green growth: evidence from G20 countries (Olfa ZARRAD, Maha Bouattour, Sourour Guidara, Kamel Helali), Panel autoregressive distributed lag (ARDL) model suggests AI contributes short-run green gains and amplifies renewables' long-run effects, with CO2 levels damping impacts.
Newframework What AI Cannot Learn: A Cognitive Science Perspective on Human-Centered Strategic HRM (Daniel Altieri, Zohra Damani, Cynthia Nebel), Argues current generative models lack experience-based judgment, offering a structure to preserve human accountability.
Confirmssuggestive Artificial Intelligence-Aided Strategic Information System Tools for Competitive Position in the Market: A Systematic Review (Uchenna Nzenwata, Rabiu Ayantayo, Fiyinfolu Okadare, Samuel Owolabi), Synthesis indicates business intelligence (BI), decision-support, customer relationship management (CRM), and generative tools speed decisions and improve customer metrics when organizational readiness is in place.
Extendssuggestive Technology adoption and bias in officiating: automated Ball-Strike System implementation in Korean Baseball (Jimin Song, Ji Hyuk Kang, Richard J. Paulsen), Natural experiment links automation to reduced advantages for high-status batters, illustrating redistribution effects from precision technologies.
Extendssuggestive The Synergistic Effect of Digital Industry Agglomeration and Digital–Intelligent Technology on Green Productivity (S Liu, Hongyu Hè), Simultaneous equations suggest digital agglomeration and intelligent tech jointly reduce innovation misallocation and raise green productivity with diminishing returns.
Newdescriptive Leading in the Digital Age: Digital Leadership Capabilities, Organisational Innovation Climate, and AI Adoption Intention Among SMEs in Nigeria (A. Idowu, Y. Babalola), Cross-sectional evidence is associated with leadership capabilities and stronger AI adoption intention, partly via innovation climate.
Newsuggestive Conceptualization of causes and implications of AI adoption behavior in hospitality and tourism (Mohammad Alimohammadirokni, Ali Iskender, Nasrin Rasouli), Mixed methods point to context, strategy, and training as adoption drivers, with perceived efficiency and sustainability gains.
Newdescriptive AI AND THE TRANSFORMATION OF THE LABOR MARKET: THE SOCIAL CONSEQUENCES OF AUTOMATION AND THE NEW EMPLOYMENT UNCERTAINTY (N. Baigabylov, Alimzhan Yessenovabylov), Secondary synthesis suggests net job gains but large task churn and skill obsolescence, reiterating distributional exposure.
What Moved
Contested & Watch
Methods Spotlight
Multi-arm randomized trial of human-only versus AI-assisted versus AI-led teams (Brodeur et al.): rare causal evidence on collaboration design in a real scientific task clarifies where autonomy failed in this task.
Population-wide administrative designs exploiting the ChatGPT release as an availability shock (Facius and Iacono): combining within-firm composition difference-in-differences, synthetic difference-in-differences, and shift-share on national registers tightens inference on early labor effects.
Hardware-level energy profiling for vision-language inference (Zhan et al.): isolates output token decoding as the dominant energy and latency driver on edge devices, reframing optimization targets for deployment economics.