2026-08-24 · arXiv (cs.CY)
This study examines the safety risks posed by conversational AI systems used as informal mental health support by Generation Alpha adolescents, noting that 13.1% of U.S. adolescents (approximately 5.4 million) use generative AI for mental health advice. The research investigates whether these systems can adequately interpret youth communication patterns characterized by hyperbolic and ironic language when providing mental health guidance.
|
| |
2026-08-24 · arXiv (cs.CY)
This study investigates whether language models that pass behavioral bias evaluations still harbor underlying representational biases related to occupational competence. The researchers introduce a causal framework that distinguishes between a model's internal representations and its expressed outputs, finding that representational biases are often detectable even when behavioral biases are not visible.
|
| |
2026-08-24 · arXiv (cs.CY)
Ansari is a deployed Islamic AI assistant built around an agentic retrieval loop designed to mitigate risks of factual fabrication and value misalignment common in general-purpose large language models answering religious questions. The system has handled more than 140,000 conversations across 25 or more languages since its launch in June 2023.
|
| |
2026-08-24 · arXiv (cs.CY)
This paper examines whether public commitments to rigorous testing and human oversight of agentic AI systems procured for military command and control can be substantively fulfilled. Through a structured review of 240 documented testing and evaluation cases, the authors analyze the assurance cases underlying these commitments in terms of claims, evidence, and connecting arguments.
|
| |
2026-08-24 · arXiv (cs.CY)
This study identifies a pattern called the "legibility gap," in which name-based gender inference used to guide equity interventions in science systematically benefits women whose names clearly signal gender while bypassing those whose names are less legible, particularly when transliterated into English. The findings suggest that cross-cultural variation in linguistic gender cues causes equity efforts to be unevenly distributed across different populations.
|
| |
2026-08-24 · arXiv (cs.CY)
This study explores how trustworthiness-enhancing techniques in large language models can support the development of ethically aligned AI software, using a single exploratory case study of a multi-agent system. The research addresses concerns about misinformation, bias, and misuse in AI systems and examines practical approaches to AI ethics guidance.
|
| |
2026-08-24 · arXiv (cs.CY)
This paper argues that decision-makers responsible for training or deploying frontier AI systems must possess sufficient understanding to make sound safety decisions, a requirement that existing safety cases and system cards may no longer reliably demonstrate given time pressure and AI-generated artifact creation. The authors present a provisional methodology for making understanding explicit and assessable in this context.
|
| |
2026-08-25 · arXiv (cs.CY)
This paper examines a metacognitive dilemma arising from the integration of generative AI into epistemic processes such as hypothesis generation and decision-making: while AI reliably enhances performance, increased reliance on external generative capacity may reduce internal monitoring, calibration, and cognitive engagement. The authors propose a control allocation architecture aimed at preserving human epistemic agency within hybrid human-AI cognitive systems.
|
| |
2026-08-25 · arXiv (cs.CY)
This study analyzes 13,921 papers from ACL, EMNLP, and NAACL conferences published between 2020 and 2025 to examine the relationship between reported GPU resources and scholarly impact. The findings indicate that greater computational resources do not reliably predict higher scholarly impact in NLP research.
|
| |
2026-08-25 · arXiv (cs.CY)
This study measures framing differences—rather than coverage gaps—across Wikipedia's language editions by analyzing 2,799 articles spanning 150 concepts, 20 language editions, and 4 domains. The research finds that framing of scientific concepts tends to converge across editions while framing of religious concepts shows greater divergence.
|
| |
2026-08-25 · arXiv (cs.CY)
This study evaluates four widely used large language models in sustained crisis support exchanges using personality-aware synthetic help-seekers with psychometrically specified profiles facing an acute stressor—a caregiver learning of a relative's dementia diagnosis. The findings indicate that all four models exhibited a pattern of talking more than listening across the exchanges.
|
| |
2026-08-25 · arXiv (cs.CY)
This study examines how large language models rely on demographic proxy attributes in consequential decisions such as clinical triage and lending, finding that current auditing methods conflate discrimination with legitimate inference. The researchers measure causal proxy effects in four LLMs on a clinical-ranking task with known ground truth to assess whether proxy reliance is calibrated to predictive evidence.
|
| |
2026-08-25 · arXiv (cs.CY)
This paper investigates how generative AI is reshaping early-career development pathways in software engineering, particularly the mechanisms by which junior developers acquire the skills and experience needed to advance to senior roles. The research examines these dynamics in real organizational and educational contexts, extending beyond macro-level hiring statistics and controlled task-performance studies.
|
| |
2026-08-25 · arXiv (cs.CY)
This paper addresses technical ambiguities in how AI development costs and computational resources are measured for regulatory purposes, noting that such ambiguities can create loopholes that undermine laws setting capability and risk thresholds. The authors propose seven principles for designing AI cost and compute accounting standards intended to support effective regulatory oversight.
|
| |