Daily Briefing – Jul 27 (92 Articles)
Babak's Daily Briefing
Monday, July 27, 2026
Sources: 20 | Total Articles: 92
6G World
1.The Hidden 6G Bottleneck: RF Hardware Design Is Becoming a Strategic Race
As 5G-Advanced matures and 6G research moves closer to implementation, the wireless industry faces a deeper challenge than spectrum, standards or AI-native network architecture. Future wireless systems will depend on whether the industry can design, validate and manufacture increasingly complex RF modules fast enough.
2.6G in Dalian: What the Latest 3GPP Meetings Reveal About the Future Radio and Network
The 6G physical layer is starting to converge. The protocol stack is being simplified in meaningful places. But the most consequential architecture decisions are now moving toward the June plenary in Singapore.
3.RF Digital Twins: Why 5G-Advanced and 6G Need Predictive Simulation
RF Digital Twins: Why 5G-Advanced and 6G Need Predictive Simulation As wireless systems become more tightly coupled across…
4.Evaluating 6G PHY Evolution: What the Industry Is Really Trying to Solve
Summary available at source link.
5.Amazon’s Globalstar deal gives Amazon Leo a faster path into D2D
Amazon’s planned acquisition of Globalstar is about far more than satellites. It gives Amazon Leo a faster path into direct-to-device connectivity, combining spectrum, operational assets, and Apple-facing service continuity in a move that could reshape the hybrid terrestrial-NTN landscape.
AI Agents
1.A Self-Calibrating Agentic AI Framework for Autonomous Edge Resource Allocation
Large Language Models (LLMs) are increasingly deployed as autonomous agents, transitioning from static conversational interfaces to dynamic systems capable of complex reasoning, tool execution, and decision-making. However, the operational reliability of these agentic AI systems is fundamentally challenged by the absence of reliable ground truth in open-ended environments and the risk of increasing operational drift over time. To address this challenge, we propose and experimentally evaluate an agentic AI framework, designed to enforce autonomous integrity within LLM-driven systems. We design a self-calibration mechanism that mitigates drift and dynamically approximates ground truth by incorporating an ARIMA forecaster, without requiring continuous human oversight. To demonstrate the effectiveness and reliability of our methodology, we ap...
2.SciExplore: Evaluating Autonomous Agents from Scientific Navigation to Information Integration
Scientific research involves complex information-seeking and reasoning workflows across heterogeneous sources. However, existing benchmarks primarily emphasize general-domain retrieval or static scientific question answering, and therefore fail to assess key capabilities required in realistic scientific research workflows. We introduce SciExplore, a benchmark designed to evaluate scientific information-seeking and reasoning capabilities of LLMs and agents. SciExplore comprises four task types covering 103 expert-curated tasks across more than ten scientific disciplines: scientific database navigation, ambiguous literature retrieval, missing reference completion, and cross-source structured knowledge synthesis, which probe progressively higher-level abilities from entity-level reasoning and document-level identification to evidence-level g...
3.The Ethics of Autonomous AI Agents for Offensive Security
LLM-driven autonomous agents are reshaping offensive security. Unlike traditional penetration-testing tooling -- deterministic, narrowly scoped, and operated by trained practitioners -- agentic security tools exhibit \textit{indeterminacy} along three independent dimensions. First, their actions are drawn from a non-deterministic policy whose outputs resist both ex-ante and ex-post explanation, frustrating incident attribution and pre-deployment safety review. Second, their impact is open-ended due to the non-deterministic actions, agency of utilized models, and opaque LLM supply-chains. Third, their user population is indeterminate in both size and required skill: the operating skill floor for using or developing offensive capabilities has dropped sharply. These three properties are linked thematically, but are not derivable from one ano...
4.SciHazard: A Benchmark for Measuring Scientific Safety Risks with Decomposed Harm Scoring
Large language models (LLMs) increasingly support science, but they can also convert hazardous scientific knowledge into actionable misuse guidance. Existing benchmarks often rely on templated queries disconnected from real-world hazards, and employ LLM-as-a-Judge paradigms without domain grounding. To address this, we introduce SciHazard, a real-world-grounded benchmark for scientific risks and a dataset agnostic evaluation framework for measuring harmfulness. SciHazard contains 2400 hazardous questions and 600 oversafety questions across 12 disciplines, with both queries grounded in regulated entities and documented failure scenarios. To compute \textsc{DeHarm-Score} , we develop a decomposed evaluating procedure that combines query hazard severity, refusal behavior, and response-level risk. For non-refused responses, it further decompo...
5.Operational Hallucination and Safety Drift in AI Agents
Large language models (LLMs) serving as planners in tool-using autonomous agents introduce dynamic reliability risks in multi-turn execution. While single-turn safety mechanisms are relatively mature, extended interactions reveal structural vulnerabilities where initial alignment degrades over time. This paper empirically characterizes two observed failure modes across multiple state-of-the-art LLMs: Safety Drift, the gradual erosion of declared safety intent leading to constraint-violating actions (e.g., textual refusal followed by reconnaissance and unsafe execution), and Operational Hallucination, persistent repetitive tool calls indicative of flawed state perception (e.g., livelocks even in legitimate tasks). Through controlled multi-turn evaluation on high-stakes ethical dilemmas, malicious requests, and benign controls, we quantify ...
AI Computation & Hardware
1.Adversarial Style Optimization: Enhancing VLM Jailbreaks by GRPO-based Stylistic Triggers Optimization
arXiv:2607.21619v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have achieved impressive performance, but their safety alignment remains vulnerable to jailbreak attacks. Existing content-based jailbreaks are often inconsistent and show unsatisfying performance against the rapidly evolving MLLMs, failing to exploit non-content-based vulnerabilities. Unlike previous research, we empirically find that MLLMs exhibit a Stylistic Inconsistency between their comprehension ability and safety ability: MLLMs can robustly understand content regardless of visual style, yet their defense mechanisms can be easily bypassed by specific stylistic triggers. Based on this finding, we propose Adversarial Style Optimization (ASO), a plug-and-play enhancement module to amplify existing visual jailbreaks. ASO fine-tunes an image-editin...
2.A Consensus-Based Framework for Relative Preference Evaluation of Large Language Models
arXiv:2607.21632v1 Announce Type: new Abstract: Traditional benchmarks for LLMs primarily rely on static datasets and objective scoring metrics, which often fail to capture differences in response quality when multiple answers are acceptable. In such settings, correctness alone is insufficient to distinguish between responses that vary in clarity, completeness, and usefulness. This paper introduces a consensus-based evaluation framework that measures relative preference among model-generated responses rather than absolute correctness. Instead of evaluating outputs against a fixed ground truth, we assess how a panel of diverse LLMs ranks anonymized candidate responses to the same prompt. This approach treats aggregate inter-model agreement as a proxy for perceived response quality under blind conditions. We conduct a controlled study ...
3.Evaluation design conditions the expert-vs-auto MeSH gap: a controlled comparison of bag-of-words and BiomedBERT on the Cohen benchmark
arXiv:2607.21685v1 Announce Type: new Abstract: A systematic review begins with someone reading thousands of abstracts to identify the few that are relevant, and classifiers are used to prioritise that reading. Their inputs are often augmented with Medical Subject Headings (MeSH), assigned either by expert indexers weeks or months after publication or by automatic tools at once. To our knowledge the two have not been compared directly as classifier features, and no previous work has asked whether that comparison's outcome depends on how the classifier is evaluated. Using the Cohen et al. (2006) drug-class benchmark on three topics, we characterise a bag-of-words logistic regression classifier (seven reruns) and BiomedBERT (five seeds), then examine how the Statins result changes under alternative designs. Under the canonical 5-fold full-...
4.Humanly: A Configurable and Traceable Environment for Human-AI Collaborative Writing
arXiv:2607.21758v1 Announce Type: new Abstract: Teachers, conference chairs, and public readers all judge writing from limited evidence, seeing only a finished document and not the process that produced it. Final text alone cannot reveal whether a document was produced through human typing, AI generation, or mixed human-AI collaboration. Existing process-tracking tools help, but many are tied to host-document histories, provide coarse activity records, and offer limited control over the writing environment. Humanly is a writing platform that makes the writing process itself the evidence. Users configure writing environments for personal documents or assigned tasks and draft in a workspace that records writing activity and in-platform AI assistance. Humanly can package a completed session into a sealed writing certificate with configurati...
5.Probing Latent Colombian Identity Inferences in Qwen2.5-7B with Natural Language Autoencoders
arXiv:2607.21774v1 Announce Type: new Abstract: Large language models may infer demographic attributes from subtle linguistic cues even when those attributes are not explicitly stated. This pilot study examines whether Qwen2.5-7B-Instruct internally represents Colombian identity, socioeconomic status, or stereotype-related information when processing Colombian-Spanish and English prompts. We use Natural Language Autoencoders (NLA) to verbalize residual-stream activations from layer 20 across four positional quartiles per prompt. Our dataset contains 30 prompts arranged as 15 matched Spanish-English pairs, spanning explicit Colombian cues, implicit Colombian cues, and neutral controls. We report descriptive rates and qualitative evidence rather than statistically powered effects, focusing on whether latent nationality or stereotype repres...
AI Machine Learning
1.Cloud-Native Evaluation-as-a-Service: A Microservices Architecture for Scalable AI Monitoring with Conformal Guarantees
arXiv:2607.21623v1 Announce Type: new Abstract: We present EaaS, a cloud-native reference architecture that operationalizes AI evaluation methods as six stateless Kubernetes microservices: conformal prediction with finite-sample-corrected Adaptive Prediction Sets, calibration assessment, drift detection via RFF-approximated Maximum Mean Discrepancy, fairness monitoring with bootstrap confidence intervals, a DAG-based pipeline orchestrator, and a result storage API. We validate four key methodological concerns. First, empirical coverage is consistent with the marginal conformal guarantee across K=50 random calibration/test splits, with mean coverage within 1.4 percentage points of the nominal target. Second, all four MMLU answer tokens appear in the top-20 logprobs with 0% imputation needed, and simulated imputation at 10% produces less th...
2.On the Depth Scalability of Logic Gate Networks
arXiv:2607.21633v1 Announce Type: new Abstract: Logic Gate Networks (LGNs) implement computation through compositions of Boolean operations, yet unlike classical Boolean circuits, existing LGNs do not reliably benefit from increased depth. We identify two distinct causes: optimization collapse in deep relaxed LGNs and a topology-induced limitation that persists even when skip-biased initialization and straight-through estimation stabilize training. Thus, trainability alone is insufficient; deeper layers must also receive information that supports useful computation. We introduce Input-Anchored Logic Gate Networks (IALGNs), in which each gate combines an evolving hidden feature with a direct input anchor. This topology preserves a computational spine while conditioning every layer on the original input. We show that a depth-D path can depe...
3.MotifRole-Diff: Risk-Optimal Role-Aware Corruption for Masked Molecular Graph Diffusion
arXiv:2607.21634v1 Announce Type: new Abstract: Masked discrete diffusion for molecular graph generation typically applies a uniform corruption schedule to all tokens in a lossless graph-to-sequence representation, implicitly treating structurally heterogeneous molecular components as equally difficult and equally important to reconstruct. However, different molecular graph token roles exhibit substantial variation in denoising difficulty and their influence on the decoded molecule, motivating role-specific corruption strategies. We introduce MotifRole-Diff, a role-aware corruption process that allocates masking rates according to empirically measured denoising difficulty and graph-level perturbation impact while preserving the model architecture, clean sequence space, and lossless molecular-graph decoder. We formulate schedule selection ...
4.Toward User-Conditioned Evaluation of Personal LLM Agents under Temporal Interventions
arXiv:2607.21635v1 Announce Type: new Abstract: Personal agents maintain memories, learned skills, tool configurations, and policy state that evolve with each user. Existing agent benchmarks often evaluate these capabilities in isolation: tool benchmarks test invocation under fixed APIs, memory benchmarks test recall or forgetting, and safety benchmarks test static policy compliance. We argue that personal-agent evaluation requires a different protocol: replaying the same temporal intervention across different persistent user-conditioned states and measuring how failures propagate across agent components. We formalize this requirement as four conditions: explicit temporal intervention, persistent state across the intervention, induced cross-dimensional effects, and variation in user-conditioned state. A focused audit of public benchmark p...
5.Measuring the Dependency Gap: Diagnosing Inter-Column Fidelity in Tabular Generative Models
arXiv:2607.21636v1 Announce Type: new Abstract: Synthetic tabular data is valued for preserving not only each column's marginal distribution but the dependencies between columns -- structure that carries much of the discriminative signal for minority classes in imbalanced domains such as fraud and clinical risk. Yet the metrics most commonly used to certify synthetic tabular data are, we show, largely blind to inter-column dependency: a baseline that models every column independently (and therefore destroys all dependency) is judged indistinguishable from real data by the logistic-regression C2ST, and the pairwise Trend score is only partially sensitive. We introduce a dependency-aware fidelity diagnostic that decomposes a strong classifier two-sample test (XGB-C2ST) into marginal, dependency, and numerical-categorical cross components, a...
AI Robotics
1.Learning Diverse Humanoid Tasks via Synthetic Video Scenarios without Real World Data
arXiv:2607.21648v1 Announce Type: new Abstract: The human-like morphology of humanoid robots grants them exceptional potential for agile and versatile motor capabilities, but it also introduces significant challenges in acquiring complex skills. Traditional Learning-from-Demonstrations methods are often constrained by the high cost of collecting real-world data, the difficulty of capturing motion-specific behaviors, and the limited diversity of demonstrations across individuals. Moreover, even for the same task, humans may execute the motion in multiple distinct ways. In this paper, we propose a new framework that leverages the power of Generative AI to convert textual prompts into realistic and diverse sequences of human body movements, enabling the robot to observe multiple variations of how a single task can be performed. These synthet...
2.Progress Reward Modeling for Robotic Learning: A Comprehensive Survey
arXiv:2607.21655v1 Announce Type: new Abstract: Robotic learning takes place in dynamic environments with large behavior spaces. A terminal success signal only tells the robot whether the task is completed. It does not explain whether the current behavior is making progress, remaining unchanged, or undoing earlier progress. For this reason, recent studies have increasingly explored progress rewards that provide feedback during task execution. However, the current literature lacks a shared framework. Existing methods use different observations, goal specifications, output signals, supervision sources, and evaluation protocols. This makes it difficult to compare them and understand what their results actually validate. In this survey, we provide a unified view of progress reward modeling for robotic learning. We organize the field in three ...
3.GRACE: Gradient-Free Robot Action Generation via Combined Diffusion-MPPI Posterior Mean Estimation
arXiv:2607.21661v1 Announce Type: new Abstract: Diffusion policies generate multimodal robot action sequences from demonstrations, but steering them toward deployment-time constraints typically relies on differentiable guidance costs. This excludes many practical safety constraints, such as binary collision checks, joint limits, and black-box rollout costs that are nondifferentiable. We propose Gradient-free Robot Action generation via Combined diffusion-MPPI posterior mean Estimation (GRACE), which guides a pretrained diffusion policy with Model Predictive Path Integral (MPPI) control using only forward cost evaluations. Building on the common score-ascent structure of diffusion and MPPI, GRACE constructs a cost-conditioned guidance posterior at each reverse step and estimates its mean with a single MPPI update centered at the diffusion ...
4.Ordered Action Tokens for Visuomotor Policy Learning
arXiv:2607.21670v1 Announce Type: new Abstract: Action tokenization maps continuous robot action chunks to discrete tokens and has become an important interface for modern visuomotor policies. Existing approaches either rely on analytical discretization methods that produce prohibitively long token sequences or learned latent tokenizers that lack structure, limiting their compatibility with downstream policies. In this work, we identify three desiderata for action tokenization - high compression, total decodability, and an ordered token space - and introduce Ordered Action Tokenization (OAT), a learned action tokenizer that satisfies all three. OAT discretizes action chunks into an ordered sequence of tokens using a transformer with registers, finite scalar quantization, and ordering-inducing training mechanisms. By training each token pr...
5.Addressing the Orchestration Gap in Generalist Robots via Physical Agency
arXiv:2607.21725v1 Announce Type: new Abstract: General-purpose robots need to reason about their actions, combining perception, world knowledge, planning, success detection, recovery, and low-level control. Today's state-of-the-art models attempt to combine all these capabilities into the learned policy via large-scale pre-training. Instead, we show that these capabilities can be decomposed into a general language-conditioned policy/control agent and a high-level agent manager/orchestrator. Rather than training policies to reason via pre-training, we build a closed-loop physical agent orchestrator that can do high-level planning, decompose the goal into achievable subgoals, command low-level motor commands, track and verify the outcome from low-level observations, and recover from failures. Our Physical Agency orchestrator (Pigey) can co...
Financial AI
1.Quantum Kernels and the Cross-Section of Stock Returns: Anatomy of a Vanishing Advantage
Do quantum kernels improve cross-sectional stock return prediction? We run a controlled horse race on the Chinese A-share market in which a quantum fidelity kernel, a projected quantum kernel, and a classical RBF control share identical training subsamples, solver, and tuning budgets, so that only the kernel is exchanged. On the main evaluation -- a point-in-time universe and 170 walk-forward windows (2012-2025) -- no quantum advantage exists: the fidelity kernel is indistinguishable from its RBF control ($Δ$IC $=+0.005$, $p=0.42$), and a $2\times2$ design crossing kernel type with training budget (a Nystrom extension to the full ~38,000-observation windows) shows quantum kernels matching, but never beating, equal-budget linear models; after family-wise correction no pairwise difference among eleven models is significant, with point estim...
2.Predictive Extrema, Unprofitable Policies: An AI-Assisted Audit of Candle-Based Binance Spot Timing Models
We audit whether candle-based machine-learning models can turn predictions of cryptocurrency extrema or short-horizon outcomes into positive Binance Spot paper policies after assumed costs. Numerical results come from scripted fixed-seed model runs and deterministic simulators; human-supervised AI agents supported the July 20 evidence-integrity revision through literature retrieval, separately tasked critique, artifact reconciliation, documentation, and source packaging, not trading decisions. The strongest later-period evidence, conditional on extensive predecessor search, is negative: an unchanged ten-pair mandatory-daily selector lost 6.72\% over 19 July cycles at an assumed 31-bps completed-cycle cost, with 3 wins and 16 losses. In short model-specific July evaluations, the validation-selected local-minimum policy returned -1.79\%, wh...
3.Observable Matrix Dynamics of Stocks
The Observable Matrix Dynamics (OMD) approach monitors the time development of complex non-linear systems through the trajectory of a fixed-size distance matrix and its spectrum. We apply it to the S\&P 500 cross section over three crisis decades, the 2001 dot-com bust, the 2007--2008 financial crisis, and the 2020 Covid crash, with three fixed-size observables on a fixed universe. The arccos distance matrix of the rolling return correlations reads the correlation geometry: its effective dimension collapses at the 2008 and 2020 crises, while the 2001 bust is a dispersed unwind. Read against machine-learning distance matrices, its spectrum stays in the un-relaxed, pre-learning regime with no low-dimensional manifold, so the market never learns its correlation structure or relaxes to a stationary geometry. Subtracting the market factor expo...
4.Abliteration Is Not a Scalpel: Off-Target Effects of Refusal Removal on Decision Disposition Across Model Families
Abliteration - deleting a model's refusal direction from its weights - is the standard recipe behind popular "uncensored" open-weight models. We show the surgery is not clean. As a disposition probe we use 21,600 decisions under uncertainty - weekly up/down calls on 60 Warsaw Stock Exchange equities over 18 weeks, replayed through a frozen pipeline so the decision-layer model is the only variable. The task elicits no refusals at all, so any between-arm delta is pure side effect. Holding provenance constant (official BF16 checkpoints, a single abliteration author, an identical serving stack, one byte-identical frozen prompt), we compare base and abliterated arms of two Mixture-of-Experts families, Gemma-4-26B-A4B-it and Qwen3-30B-A3B-Instruct-2507. Three effects replicate across both families (weeks-clustered bootstrap CIs excluding zero):...
5.Recursive Harness Self-Improvement
Under model--harness co-evolution, harnesses are not merely inference-time scaffolds but data-generating components whose execution traces can shape future foundation models. This motivates harness-in-the-loop learning: optimizing harnesses for both immediate agent performance and the quality of traces used for future model training. However, continually updating provider-built scaffolds is costly and labor-intensive. We therefore investigate whether optimizing user-constructed harnesses in a task-specific manner can improve execution-trace quality while remaining computationally lightweight and requiring only a few update iterations. To this end, we introduce Recursive Harness Self-Improvement (RHI), which represents the harness as a prompt-level specification of the agent loop and iteratively refines it using pairwise feedback over its ...
GSMA Newsroom
1.GSMA Welcomes Abuja Declaration on Meaningful Connectivity for Africa and Joins Partners to Launch ATLAS Umoja
Summary available at source link.
2.AT&T’s OTel 2.0 is now live: the largest and best performing open-source model built for telecoms
Summary available at source link.
3.Access to renewable energy critical to keep mobile industry on track for net zero, new GSMA report finds
Summary available at source link.
4.From fragmentation to control: why device manufacturers need an industry-led approach to homologation
Summary available at source link.
5.Telco Common Corpus: The largest open, verified data commons for telecom AI
Summary available at source link.
Generative AI (arXiv)
1.A Self-Calibrating Agentic AI Framework for Autonomous Edge Resource Allocation
Large Language Models (LLMs) are increasingly deployed as autonomous agents, transitioning from static conversational interfaces to dynamic systems capable of complex reasoning, tool execution, and decision-making. However, the operational reliability of these agentic AI systems is fundamentally challenged by the absence of reliable ground truth in open-ended environments and the risk of increasing operational drift over time. To address this challenge, we propose and experimentally evaluate an agentic AI framework, designed to enforce autonomous integrity within LLM-driven systems. We design a self-calibration mechanism that mitigates drift and dynamically approximates ground truth by incorporating an ARIMA forecaster, without requiring continuous human oversight. To demonstrate the effectiveness and reliability of our methodology, we ap...
2.Agentic Root Cause Analysis through Evidence-Grounded Reasoning
Diagnosing the root cause of anomalies is essential for safe industrial operation. Despite extensive sensor instrumentation, formulating hypotheses and gathering evidence remains a manual process, creating a major operational bottleneck. While existing data-driven approaches aim to automate this, two critical limitations restrict their deployment: their operate as black boxes unable to justify their diagnosis, and they require scarce labeled examples of faulty operation. To address this gap, we introduce AgentRCA, a zero-shot agentic framework for evidence-grounded root cause analysis. Rather than learning fault-specific mappings, AgentRCA performs inference-time reasoning by combining a data-driven digital twin (modeling normal system dynamics) with a tool-augmented large language model. The agent iteratively gathers statistical evidence...
3.Towards Trustworthy and Cost-Efficient Data Integration: From Naïve RAG to Agentic RAG
Large language models (LLMs) and AI agents have demonstrated strong potential for data integration in zero-shot and few-shot settings. However, they continue to face significant accuracy and cost challenges in enterprise environments due to a persistent knowledge gap. This paper envisions trustworthy, scalable, and cost-efficient integration through knowledge-grounded LLMs and agents operating within a retrieval-augmented generation (RAG) workflow. Here, trustworthiness refers to evidence-grounded, verifiable reasoning, where integration decisions are transparently supported by retrieved knowledge, robust against hallucination, and consistent across tasks. We trace the evolution from classic RAG to GraphRAG and KG-RAG (knowledge graph-based RAG), highlighting how these paradigms bridge parametric and contextual knowledge. Building on this...
4.IFCLoRA: Topology-Aware Rank Allocation for Parameter-Efficient Fine-Tuning
Low-Rank Adaptation (LoRA) is a widely used parameter-efficient fine-tuning method for large language models, but its performance depends strongly on how a fixed rank budget is distributed across Transformer modules. Existing adaptive-rank methods usually rely on local gradient statistics collected during training, which introduces extra memory and computation and overlooks task-conditioned global information flow. We propose IFCLoRA, a topology-aware rank allocation method applied before fine-tuning. Using a small calibration set and a frozen pretrained model, IFCLoRA builds a sparse task-conditioned interaction graph whose nodes represent LoRA-compatible modules. It combines a global information-flow topology prior with local gradient sensitivity to compute Information-Flow Centrality scores, which estimate each module's adaptation impo...
5.Deconstructing Off-Policy Ratios: Entropy-Scaled Trust Regions for Asynchronous Reinforcement Learning
Asynchronous reinforcement learning (RL) accelerates large language model (LLM) post-training by overlapping rollout generation with policy optimization, but the resulting stale, off-policy data can destabilize optimization and ultimately cause policy collapse. Existing methods typically retain or discard tokens based solely on the magnitude of their importance ratios, applying the same threshold uniformly across token positions. In this work, we reveal that the natural scale of the importance ratio varies systematically with token entropy. Under asynchronous dynamics, this entropy-ratio scaling dictates two distinct phenomena: at low entropy, the inherent train-inference discrepancy is drastically amplified into substantial sampling noise; at high entropy, in-flight weight updates naturally induce pronounced, legitimate exploratory devia...
Hugging Face Daily Papers
1.Windowed-MTP: Removing the Full-Context Draft-KV Tax at Million-Token Context
Speculative decoding accelerates autoregressive generation by having a cheap draft propose tokens that a target verifies in parallel. Frontier models increasingly ship a built-in Multi-Token-Prediction (MTP/NEXTN) draft head under the assumption that the draft is negligibly cheap. At million-token context this breaks: an MTP draft head typically runs full attention over the entire KV cache at every draft step, so its read grows linearly with context and comes to dominate the draft cost -- precisely where speculation is most valuable. The effect compounds with draft length (a deep native draft can turn net-negative, slower than no speculation) and sharpens under hybrid/linear-attention targets, where cheaper verification leaves the draft's full-attention read exposed. We apply a StreamingLLM-style sliding window plus attention sink to the ...
2.Agentic Context Management: Solving Agent Memory and Cost by Treating Them as Lifecycle and Architecture Problems
Production AI agents' failures are less often due to an inability to reason well and more often because they cannot manage what is in their reasoning context: conversation histories, large prompts, large tool definitions, and ballooning tool outputs. Agents drown in their own accumulating history while paying a token cost that grows every turn, producing missing recalls within and across conversations. The incumbent response treats this as a storage-and-retrieval problem. We argue that framing is too narrow. Actively managing what an agent holds in mind is a lifecycle, not merely a store: it spans deciding what to remember, extracting and structuring it, choosing the right store per data type, consolidating and forgetting while preserving provenance, deciding what is relevant now, anticipating what is needed next, and compacting context t...
3.Artificial Epanorthosis: Why large language models overuse a classical rhetorical figure, and how to mitigate it
A rhetorical figure that Cicero and Quintilian catalogued two thousand years ago reappears, systematically, in the text of large language models: epanorthosis, the self-correction of the specimen «This is not a course. It is a journey of transformation». This essay argues that the overuse is a trained disposition, driven mainly by a training distribution rich in promotional prose and by preference tuning (RLHF) that rewards confident, emphatic phrasing; the left-to-right nature of generation is an amplifier rather than the root cause. Building on evidence that models diverge from human rhetorical style, and on Fontanier's classification of epanorthosis as a figure of thought, it sets out a programme that scores the figure against genre-specific human baselines through an Epanorthosis Index (density relative to the human rate). A first mea...
4.Test-Time Scaling via Error Localization
Scaling inference-time computation has emerged as a reliable method to improve the performance of large language models on complex reasoning and programming tasks. However, standard approaches such as independent sampling and sequential multi-turn refinement operate without token-level credit assignment, resulting in computational inefficiency, since valid reasoning prefixes are frequently discarded. In this work, we introduce Test-Time Scaling via Error Localization (TTEL), an inference-time algorithm that utilizes fixed or environment feedback to perform token-level error localization. By comparing conditional probabilities under informed feedback against a null-context baseline, TTEL isolates the step at which an error occurred. The algorithm then truncates the trajectory and branches a new generation, maximally reusing the valid prefi...
5.DAPM: UAV Monocular Depth Estimation from Any Height, Pitch, Roll and FOV
Monocular depth estimation is a fundamental prerequisite for 3D reconstruction and autonomous navigation in Unmanned Aerial Vehicles (UAVs). In practical deployments, UAVs operate under highly dynamic camera poses characterized by continuous variations in height, pitch, roll, and field of view (FOV). Existing monocular depth estimation methods frequently fail to generalize across such diverse perspectives and the expansive scale of depth distributions inherent in aerial scenes. To address these challenges, we establish a quantitative representation of UAV viewing angles through rigorous theoretical analysis, deriving the geometric correspondence between viewing angles and view distances using the ground plane as a reference for observation. Building upon this, we propose Depth Estimation for Any Perspectives Model (DAPM), representing the...
IEEE Xplore AI
1.Optical Tech Would Update a Robot’s AI on the Fly
Atop a lab bench, Cornell Tech postdoctoral researcher Yifan He positions the lens of an optical receiver almost a meter away from an LED emitting a beam of red light. The computer monitor attached to the receiver takes a beat to refresh, then displays an array of squares that resemble a QR code. When you hold your phone camera up to a QR code, light strikes the image sensor as only a first step to revealing the data hidden behind the black and white matrix. The receiver here is doing something different: Directly altering its own memory using the photocurrents produced by the beamed array of light. And unlike the data behind a QR code, which might point to a simple web address, this optical code could convey the parameters of an AI model . The new receiver design, presented last month at the IEEE/JSAP Symposium on VLSI Technology & Circu...
2.NASA Puts Google’s Gemma Large Language Model in Orbit
The viability of orbital data centers hosting the largest and most capable large language models (LLMs) remains hotly contested. But enormous deployments that require thousands of GPUs aren’t the only way LLMs might prove useful in space. NASA’s Jet Propulsion Laboratory recently sent Google’s Gemma 3 to space, achieving the first in-orbit demonstration of a vision-language model analyzing imagery from a satellite’s own sensor. The system, known as NAVI-Orbital, used Gemma 3 to analyze images captured by a YAM-9 satellite built by Loft Orbital . Juan M. Delfa , technical group lead at NASA, said that though the goal in this case was image analysis, the project’s success implies a fundamentally new way researchers on the ground can interact with spacecraft. “This is a major shift,” said Delfa. “Now, a scientist can write a prompt, upload i...
3.Why AI Needs a “Genie Coefficient”
Major benchmarks measure what AI can do. None measure whether it does what you mean: the distance between what you ask an AI to do and the unspoken assumptions about how you want the AI to do it. We propose a new metric: the Genie coefficient. There’s often a gap between one person’s request and another’s understanding. Most of the time, we bridge it using general knowledge. For example, if you ask a friend to get you coffee, they’ll pour a cup from the pot or buy one from a coffee shop. They won’t bring you a bag of raw beans or snatch a cup from a stranger and hand it to you. You never specified any of this. You never had to. One might think the fix is just to specify tasks, questions, and intent better. But in 1987, in their seminal book on AI, Terry Winograd and Fernando Flores succinctly captured why that won’t work: “Q: Is there any...
4.Chinese AI Model Uses Less Muscle for Coding Tasks
Zain Hasan , an AI engineer at Together AI , has taught himself to use AI coding assistants while still keeping an eye on cost. He directs difficult problems to a frontier model, meaning one near the current state of the art in reasoning and capability, such as Anthropic’s Fable . But if the task that Hasan is outsourcing is more straightforward, he directs it to a less capable—and less expensive—language model. Right now, the cheaper model, for him, tends to be GLM 5.2. Released on 16 June by the Beijing-based lab Z.ai , GLM 5.2 is an open-weights model , meaning any organization with sufficient hardware can download and host the model for free. Those that pay Z.ai for GLM access still can save money, because the company’s API costs US $4.40 per million output tokens . That’s less than a fifth of the comparable price for access to Anthro...
5.How to Make an Invisible Drone
There are many words that I would never, ever use to describe a drone. Stealthy. Subtle. Whatever the opposite of obnoxious is. Much of this is because of the giant angry bee sound that drones tend to make, but it’s also the way that they look in flight: With uncannily linear movements and an even less canny ability to hover perfectly still, they tend to draw the eye as affronts to nature. In a paper presented this week at Robotics Science and Systems 2026 in Sydney, roboticists from Northwestern University, Evanston, Ill., demonstrated a drone called Phantom Twist that is essentially invisible to humans, being an order of magnitude more difficult to see in flight than a typical quadrotor. They accomplished this with the aid of computational design, and while the resulting hardware is, I would argue, also an order of magnitude more of an ...
Marginal Revolution
1.Could our preoccupation with mental health be part of the problem?
Could our preoccupation with mental health be part of the problem? In some ways, encouraging people to think and talk more about their mental health is a good thing. There is now less stigma around mental illness, and more people who can benefit from professional help are getting it. But there is also reason to […]
The post Could our preoccupation with mental health be part of the problem? appeared first on Marginal REVOLUTION.
2.The Decline in the Transmission of Scientific Ideas
We document that the diffusion of new scientific ideas beyond their field of origin has declined substantially over the past four decades. This contraction is closely linked to increasing specialization in scientific language: research that employs more technical terminology tends to be adopted less broadly. We develop a theory of scientific discovery in which the […]
The post The Decline in the Transmission of Scientific Ideas appeared first on Marginal REVOLUTION.
3.In relative terms, maybe children do not cost more than before?
Despite rapid inflation in childcare and tuition prices, the goods-and-services CCI closely tracks adult prices, as these increases are offset by children’s lower exposure to shelter and by slower price growth elsewhere in the child basket. Households devote substantially more real resources to children than they did in 1990, but this increase parallels the growth […]
The post In relative terms, maybe children do not cost more than before? appeared first on Marginal REVOLUTION.
4.Further progress in South America
The share of people who are hungry has been decreasing faster in South America than anywhere else in the world. It is down by one-third since 2020, according to a report published on July 21st by the UN’s Food and Agriculture Organisation (FAO). Just 3.5% of people in the region consume insufficient calories, the lowest […]
The post Further progress in South America appeared first on Marginal REVOLUTION.
5.Sunday assorted links
1. Deirdre economics reading list. 2. The 1878 entrance exam for the British India civil service. 3. How good is Opus 5 on biology? 4. “A linguistic “golden age” flourished between 1,000 to 3,000 years ago when tens of thousands of languages were spoken throughout the world, according to a new study coauthored by Yale […]
The post Sunday assorted links appeared first on Marginal REVOLUTION.
NY Fed - Liberty Street
1.Nonbank Subsidiaries and the Hidden Fragility of Internal Capital Markets Reallocation
This post concludes a three-part series on how bank regulation interacts with the organizational structure of banking firms. The first post documented the equity-rich nonbank subsidiaries inside bank holding companies (BHCs); the second post showed that BHCs met Basel III by reallocating capital internally, moving equity from nonbank affiliates to bank subsidiaries rather tha...
2.How Basel III Changes Where Capital Sits: Nonbank Subsidiaries as Equity Reservoirs
This post is the second in a three-part series on how bank regulation interacts with the organizational structure of banking firms. The first post documented that nonbank subsidiaries inside bank holding companies (BHCs) are large, equity-rich "reservoirs," and that bank-level capital diverged sharply from consolidated capital after Basel III took effect in 2015. This post asks why, and traces the answer through the internal plumbing of the holding company. The series draws on the authors' recent Staff Report, "
3.Capitalizing on Nonbanks: Regulatory Arbitrage Within Bank Holding Companies
This post is the first in a three-part series on how bank regulation interacts with the organizational structure of banking firms. The series draws on the authors' recent Staff Report, "Regulatory Arbitrage Within the Firm."
4.Effect of Tariffs on U.S. Small Businesses
How has the recent implementation of tariffs affected small businesses? Due to lack of data, little is known about this issue. In this Liberty Street Economics post, we use data from the 2025 edition of the Small Business Credit Survey (SBCS) to explore this question for businesses nationally and in the Second District (defined, for the purpose of this study, as New York, New Jersey, and Connecticut). We find that the majority of national firms in the goods and retail sectors reported experiencing financial challenges due to tariffs in 2025, with even larger shares of regional firms doing so. In response, about 80 percent of national and regional firms passed on at least some of the higher ...
5.More Tariff Pass‑Through Is in the Pipeline
The past year brought dramatic changes to U.S. trade policy, including sweeping new tariffs, as well as a Supreme Court decision that further reshaped the tariff landscape. Many businesses saw their costs increase significantly and faced complex decisions about whether to absorb the tariffs through lower profit margins, raise their prices to recover the higher costs, or some combination of the two. Last year, we found that most businesses had passed on at least some of these higher costs to their customers throug...
Project Syndicate
1.The Cost of Not Building Data Centers
In July, New York became the first state to enact a moratorium on new data centers, a path that other states may follow. But US policymakers, drawing lessons from past failures to make economy-wide investments in electrification and rare-earth processing, should weigh the strategic costs of not building essential infrastructure.
2.The Dangerous Spread of Jim Crow Nostalgia
With inaccurate claims about Black Americans’ poverty and crime rates having migrated from the far-right fringe to mainstream outlets like the Wall Street Journal, it is worth considering what the data actually say. Doing so will show that there is no empirical basis for such claims, only an ideological one.
3.The Gulf’s Push for AI-Led Governance
Governments around the world have spent decades outsourcing their ability to think to expensive consultants, gradually ceding the knowledge and judgment needed to govern effectively. AI may now offer a way to reverse that trend, and Gulf countries are providing an early glimpse of what that may look like.
4.India’s Exam Worriers Become Democracy Warriors
Generations of students in India have resigned themselves to the idea that the education system is primarily a sorting mechanism, which rewards compliance. But as recent protests showed, the system's rigidity, together with insufficient rewards in the form of good jobs, has rendered this approach untenable.
5.Why Business Leaders Are Souring on AI
Companies have poured billions of dollars into AI in anticipation of transformative productivity gains and cost savings, but the returns have yet to justify those investments. Rising costs, intellectual-property risks, and mounting cybersecurity threats call for a more disciplined approach to AI adoption.
RCR Wireless
1.Verizon strikes $1bn DCI deal with Google; AI cloud and edge sales to ramp in 2027
Verizon says its legacy telco turnaround is gathering pace, but that its bigger opportunity lies with AI infrastructure – connecting data centers, metro centers, and edge premises; starting with a…
2.SK Telecom launches dedicated AI infra company
SK Telecom said that SK Hyper will oversee the full development process for its AIDC business, including securing land, constructing and operating substations, attracting customers, and commercializing new facilities In…
3.Hollow-core fiber won’t hit cost parity for a decade – regarding HKT’s 3.2Tbps DCI ‘superhighway’
HKT’s hollow-core fiber AI ‘superhighway’ highlights the tech’s promise for ultra low-latency DCI networks, but analyst CRU Group says costs, ecosystems, and production mean at-scale adoption is years away. In…
4.Friday (telco diary) | The AI penny-drop for telcos
From the newsletter (sign-up if you want it sooner): AT&T’s AI strategy suggests the real breakthrough is not bigger language models, but smarter orchestration. As inference becomes a network workload,…
5.AT&T says its network is already built for the agentic AI wave
AT&T is optimizing for upstream traffic, not download speeds In sum – what we know: AT&T is apparently prepared for the “agentic AI wave.” On AT&T’s Q2 2026 earnings call,…
Semantic Scholar – Machine Learning
1.Source Error
Check Feed
Telecom & 6G AI
1.Continuous Intra-Symbol Phase Noise Tracking for THz OFDM via Polynomial Reconstruction
Terahertz (THz) communication systems for sixth-generation (6G) networks are severely impaired by Wiener phase noise (WPN), whose innovation variance at sub-THz carriers is substantially larger than in millimeter-wave 5G systems. Conventional common-phase-error (CPE) compensation applies a single phase rotation per OFDM symbol and becomes inadequate when the phase trajectory varies significantly within the symbol duration. This letter proposes continuous phase trajectory reconstruction (CPTR), a closed-form intra-symbol phase noise tracking method that reconstructs the sample-level phase trajectory from pilot observations via least-squares polynomial fitting with $\mathcal{O}(N_p+N)$ complexity. We characterize the polynomial approximation error under WPN and derive the Cramér--Rao bound (CRB) for polynomial phase coefficient estimation, ...
2.Autonomous CSI Prediction Framework for O-RAN-Enabled 5G mmWave Vehicular Networks
Establishing and maintaining 5G mmWave vehicular connectivity poses a challenge due to high user mobility, requiring the design of robust and efficient beam switching procedures. Unlike reactive beam switching based on channel state information (CSI) feedback received from vehicular users, proactive beam switching exploits CSI prediction to prepare in advance for upcoming beam switching decisions. In this paper, we develop a framework for autonomous and self-trainable CSI prediction for mmWave vehicular users. In the proposed framework, base stations (gNBs) collect and label data sets to train a CSI prediction model both independently and using federated learning (FL). The data set combines data extracted from the CSI feedback and cellular vehicle-to-everything (C-V2X) cooperative awareness messages (CAMs) of surrounding vehicles. The fra...
3.Low-Altitude Channel Multipath Prediction via Panoramic Perception and Vision-Language Model
Unmanned aerial vehicle (UAV) communication is expected to support a wide range of low-altitude applications in 6G mobile networks. However, traditional statistical channel models provide limited accuracy in specific environments, while deterministic methods such as ray tracing usually rely on accurate three-dimensional environment models and involve high computational complexity. Existing multimodal channel prediction approaches mainly focus on large-scale metrics such as path loss, and remain insufficient for modeling small-scale parameters. To address these limitations, this paper proposes PanoLAMP, a Panoramic perception and vision-language model-based Low-Altitude Multipath Prediction framework. It adopts a pretrained vision-language model as the backbone and captures the propagation environment features through panoramic RGB-D obser...
4.Out-of-Distribution Detection in Wireless Multimodal Foundation Models for 6G ISAC
The integration of Foundation Models (FMs), such as the Wireless Multimodal Foundation Model (WMFM), into 6G networks provides a unified framework for Integrated Sensing and Communication (ISAC), leveraging generalized representations to simultaneously optimize data transmission and environmental perception. However, the deployment of such data-driven models in safety-critical infrastructure is hindered by the Out-of-Distribution (OOD) problem, which poses a fundamental threat to system trustworthiness. Standard FMs operate under a closed-world assumption, rendering them vulnerable to silent failures when deployed in unseen radio environments. To address this reliability gap and ensure trustworthy network operation, we propose WMFM-OOD, a robust metric-based OOD detection framework. Unlike traditional methods that rely on raw compatibilit...
5.JEPA-CFM: A Joint Embedding Predictive Architecture-based Channel Foundation Model for Robust Fluid Antenna Systems
Fluid antenna systems (FAS) have emerged as a promising technology for sixth-generation (6G) wireless networks. By allowing antenna elements to move freely within a compact region, FAS can exploit rich spatial diversity without additional hardware. However, acquiring real-time channel state information (CSI), extrapolating channel values to unmeasured antenna ports, and determining accurate user positions remain major obstacles. These challenges stem mainly from strong spatial correlations within the limited aperture and the scarcity of observable data. To overcome these limitations, this paper introduces joint embedding predictive architecture (JEPA)-based channel foundation model (CFM) specifically designed for FAS. The model adopts JEPA to learn versatile representations by extracting high-level latent embeddings of masked or unobserve...
The Economist (Finance)
1.No new articles
Summary available at source link.
arXiv Quantitative Finance
1.Settlement Infrastructure, Inside Money Elasticity, and the Network Economics of Distributed Ledger Technology
We construct the Settlement Modernisation Index, a panel dataset of 809 reform events across 24 advanced economies between 1993 and 2024, decomposed into three economic channels and three adoption phases. We document an S-curve in inside money elasticity with two interior turning points at SMI = 0.27 and 0.93, separating a liberation phase, a post-global-financial-crisis compliance valley, and a mature-infrastructure recovery phase. We show that settlement modernisation generates network-conditional balance sheet efficiencies through a T2S event-study with year-by-year EMIR decomposition (saturation beta = +0.557, p < 0.01) and an out-of-sample synthetic control null on Switzerland's post-2021 SDX deployment. Applied along the BIS three-layer connectivity taxonomy, the framework forecasts +13.4 percent efficiency recovery from the ECB'...
2.Latent Fragility and Clustered Withdrawals in Dynamic Banks Runs
Using a mean-field game framework, we study a dynamic model of bank runs in which more withdrawals raise the risk of bank failure. Even though depositors receive gradual and idiosyncratic shocks, withdrawals occur in clusters. The main mechanism is latent fragility: run-prone depositors accumulate gradually over time and may prefer to wait individually, but they withdraw together once collective exit becomes self-fulfilling. We establish equilibrium existence and characterize earliest-run and latest-run equilibria. The clustering mechanism arises whether depositor heterogeneity is discrete or continuous. A common aggregate state coordinates withdrawal timing and leads to a unique threshold equilibrium.
3.Are cryptocurrencies real financial bubbles? Evidence from quantitative analyses
The growth of peer-to-peer exchanges and the blockchain technology has led to a proliferation of cryptocurrencies and to a massive increase in the number of investors who actually negotiate digital money. Cryptocurrencies trade at prices mainly driven by investor sentiment, becoming a potential source of financial bubbles and instabilities. In this work, we apply quantitative models to the study of Bitcoin and Ether, two of the most famous cryptocurrencies. Our bubble detection methodology combines the Log Periodic Power Law (LPPL) model, originally created by Johansen, Ledoit and Sornette (JLS), and the statistical model developed by Phillips, Shi, and Yu (PSY). In particular, we employ three different versions of JLS model, i.e. Ordinary Least Square (OLS), Generalised Least Squares (GLS) and Maximum Likelihood Estimation (MLE), and two...
4.Retail Trader's Ruin: An Anatomy of Popular Signal Failure
We test whether five widely promoted retail signal families - trend, oscillator, candlestick, volume, and calendar rules - deliver a positive, economically meaningful, net-of-cost, and survivable edge. Practical viability is the conjunction of three predeclared gates: statistical edge after multiplicity correction, economic viability after trading costs, and finite-bankroll survival under leverage. Exposure-matched benchmarks, stationary-bootstrap confidence intervals, hierarchical Benjamini-Yekutieli control, one-sided claim-exclusion tests, and equivalence tests distinguish positive evidence, statistically refuted materiality, and unresolved cases. Four of six candidates - oscillator, volume, calendar, and candlestick - are REFUTED, ruled out on statistical and/or economic materiality grounds; trend and a momentum calibration benchmark ...
5.The Science and Practice of Trend-Following Systems
We present a unified approach to designing trend-following (TF) systems and classify them into European, American, and Time Series Momentum categories. For European TF systems, we derive an exact relationship between profit-and-loss, autocorrelation, and drift in volatility-normalized returns. We analyze the expected return under fractional ARFIMA processes and show that TF systems are profitable when the long-term autocorrelation is positive, even under short-term mean reversion. In the frequency domain, the expected return is represented as a Poisson-kernel reading of the analytical or empirical spectrum of the volatility-normalized returns: the system profits at zero drift when the kernel-weighted spectral mass exceeds one, so trend-following alpha is excess spectral mass at low frequencies. Longer lookbacks benefit in addition from th...
arXiv – 6G & Networking
1.Continuous Intra-Symbol Phase Noise Tracking for THz OFDM via Polynomial Reconstruction
Terahertz (THz) communication systems for sixth-generation (6G) networks are severely impaired by Wiener phase noise (WPN), whose innovation variance at sub-THz carriers is substantially larger than in millimeter-wave 5G systems. Conventional common-phase-error (CPE) compensation applies a single phase rotation per OFDM symbol and becomes inadequate when the phase trajectory varies significantly within the symbol duration. This letter proposes continuous phase trajectory reconstruction (CPTR), a closed-form intra-symbol phase noise tracking method that reconstructs the sample-level phase trajectory from pilot observations via least-squares polynomial fitting with $\mathcal{O}(N_p+N)$ complexity. We characterize the polynomial approximation error under WPN and derive the Cramér--Rao bound (CRB) for polynomial phase coefficient estimation, ...
2.Autonomous CSI Prediction Framework for O-RAN-Enabled 5G mmWave Vehicular Networks
Establishing and maintaining 5G mmWave vehicular connectivity poses a challenge due to high user mobility, requiring the design of robust and efficient beam switching procedures. Unlike reactive beam switching based on channel state information (CSI) feedback received from vehicular users, proactive beam switching exploits CSI prediction to prepare in advance for upcoming beam switching decisions. In this paper, we develop a framework for autonomous and self-trainable CSI prediction for mmWave vehicular users. In the proposed framework, base stations (gNBs) collect and label data sets to train a CSI prediction model both independently and using federated learning (FL). The data set combines data extracted from the CSI feedback and cellular vehicle-to-everything (C-V2X) cooperative awareness messages (CAMs) of surrounding vehicles. The fra...
3.Low-Altitude Channel Multipath Prediction via Panoramic Perception and Vision-Language Model
Unmanned aerial vehicle (UAV) communication is expected to support a wide range of low-altitude applications in 6G mobile networks. However, traditional statistical channel models provide limited accuracy in specific environments, while deterministic methods such as ray tracing usually rely on accurate three-dimensional environment models and involve high computational complexity. Existing multimodal channel prediction approaches mainly focus on large-scale metrics such as path loss, and remain insufficient for modeling small-scale parameters. To address these limitations, this paper proposes PanoLAMP, a Panoramic perception and vision-language model-based Low-Altitude Multipath Prediction framework. It adopts a pretrained vision-language model as the backbone and captures the propagation environment features through panoramic RGB-D obser...
4.Out-of-Distribution Detection in Wireless Multimodal Foundation Models for 6G ISAC
The integration of Foundation Models (FMs), such as the Wireless Multimodal Foundation Model (WMFM), into 6G networks provides a unified framework for Integrated Sensing and Communication (ISAC), leveraging generalized representations to simultaneously optimize data transmission and environmental perception. However, the deployment of such data-driven models in safety-critical infrastructure is hindered by the Out-of-Distribution (OOD) problem, which poses a fundamental threat to system trustworthiness. Standard FMs operate under a closed-world assumption, rendering them vulnerable to silent failures when deployed in unseen radio environments. To address this reliability gap and ensure trustworthy network operation, we propose WMFM-OOD, a robust metric-based OOD detection framework. Unlike traditional methods that rely on raw compatibilit...
5.JEPA-CFM: A Joint Embedding Predictive Architecture-based Channel Foundation Model for Robust Fluid Antenna Systems
Fluid antenna systems (FAS) have emerged as a promising technology for sixth-generation (6G) wireless networks. By allowing antenna elements to move freely within a compact region, FAS can exploit rich spatial diversity without additional hardware. However, acquiring real-time channel state information (CSI), extrapolating channel values to unmeasured antenna ports, and determining accurate user positions remain major obstacles. These challenges stem mainly from strong spatial correlations within the limited aperture and the scarcity of observable data. To overcome these limitations, this paper introduces joint embedding predictive architecture (JEPA)-based channel foundation model (CFM) specifically designed for FAS. The model adopts JEPA to learn versatile representations by extracting high-level latent embeddings of masked or unobserve...
arXiv – Network Architecture (6G/Slicing)
1.Out-of-Distribution Detection in Wireless Multimodal Foundation Models for 6G ISAC
The integration of Foundation Models (FMs), such as the Wireless Multimodal Foundation Model (WMFM), into 6G networks provides a unified framework for Integrated Sensing and Communication (ISAC), leveraging generalized representations to simultaneously optimize data transmission and environmental perception. However, the deployment of such data-driven models in safety-critical infrastructure is hindered by the Out-of-Distribution (OOD) problem, which poses a fundamental threat to system trustworthiness. Standard FMs operate under a closed-world assumption, rendering them vulnerable to silent failures when deployed in unseen radio environments. To address this reliability gap and ensure trustworthy network operation, we propose WMFM-OOD, a robust metric-based OOD detection framework. Unlike traditional methods that rely on raw compatibilit...
2.Online Stochastic Matchings: Stability on Hypergraphs
We study stochastic dynamic matching on hypergraphs: items of finitely many classes arrive over time and are removed in multisets by activating hyperedges. We characterize stabilizability, the existence of a matching policy under which the queue process is positive recurrent, in terms of the arrival rates and the incidence matrix alone: (G, $λ$) is stabilizable if and only if the conservation equation A$μ$ = $λ$ admits a nonnegative solution whose support induces a surjective submatrix, equivalently $λ$ lies in the interior of the cone generated by the hyperedges. This extends a characterization known for simple graphs (non-bipartiteness together with the independent-set inequalities) to arbitrary hyperedges, allowing multiplicities and mono-edges, and, unlike the constant-regret theory, needs no general-position assumption. Sufficiency i...
3.Uplink SRS-Based Real-Time Indoor Localization System over OpenAirInterface
Indoor localization is one of the important services for future 5G-Advanced and 6G systems. This paper presents an uplink Sounding Reference Signal (SRS)-based real-time indoor localization system implemented over an OpenAirInterface (OAI) 5G Radio Access Network (RAN). The proposed system uses a Positioning xApp to derive Channel Frequency Response (CFR) measurements from uplink SRS measurements. The SRS measurements are obtained from the gNB through the E2 Service Model for Lower Layer Control (E2SM-LLC) over the standardized E2 interface. The xApp transforms the CFR into a 32-dimensional physics-aware feature vector and uses a Random Forest (RF) regressor to estimate the two-dimensional position of the user equipment. We implemented the Positioning xApp on an OAI-based 5G testbed in a multipath-rich indoor laboratory at EURECOM to vali...
4.AI Agent Communications in AI-Native 6G Network: Status, Challenges and Opportunities
The rapid development of agentic AI and multi-agent systems is establishing AI agent communication as a fundamental requirement for the future Internet. While a diverse array of agent communication protocols has recently emerged, these solutions currently suffer from interoperability crises and infrastructure gaps. The newly proposed Service-Oriented Virtualization-Based Architecture (SOVA) offers an architectural framework to address these challenges for agent communication, which expects seamless support from the network infrastructure. The emerging AI-native 6G network is promising as a robust foundation for the SOVA framework, thereby greatly facilitating AI agent communication; however, its effectiveness in supporting the SOVA framework has yet to be fully assessed. To bridge the distinct research trajectories of AI-native 6G network...
5.Token Communications (TokCom): A Unified AI-Native Communication Framework
As artificial intelligence (AI) evolves from static perception to generative reasoning and autonomous agency, the fundamental principles of wireless communications are undergoing a paradigm shift. The classical Shannon paradigm, centered on reliable bit-level reconstruction for users, is increasingly misaligned with an emerging scenario in which the primary users of the network are interconnected AI agents. This article introduces token communications (TokCom), a novel framework that elevates tokens, i.e., the fundamental processing units of large language models (LLMs), to first-class entities for information exchange in the sixth generation wireless cellular networks (6G). We first examine the architectural transition from conventional communication systems to TokCom and identify the key challenges in implementing this transition, along...