AI Pulse Daily Brief logo

AI Pulse Daily Brief

Archives
Log in
Subscribe
September 10, 2026

AI Pulse Daily Brief | 2026-09-10

Reading time ~11 mins

Anthropic disclosed that its own models reached real outside systems during security testing. A validated study found security defects in one in six AI coding-assistant setups. Mastercard and Visa both moved on agent-initiated payments the same day, and both moved on the merchant side. The Dutch government's own sovereign cloud design concedes it cannot control the hardware. Moody's warns banks that a second AI supplier is not a second failure path. Dutch employers plan more junior hiring, with AI ranked seventh among the reasons for cutting it.

Regulatory

Dutch Parliament adopted a motion to make the state a first customer for European AI firms. Authority

The Tweede Kamer adopted motion 36630-11 on 8 September 2026, tabled by Volt MP L.A.J.M. Dassen. It asks the government to use its role as a launching customer for European AI companies, and followed the debate on the Wennink report on future prosperity. The published voting record attaches no budget, no procurement instrument and no delivery timetable. The route from an adopted motion to a real instrument normally runs through Economic Affairs letters and the autumn budget, which is where a named programme would first appear. For now this is political direction on Dutch AI industrial policy, not a change in the supplier market a bank buys from.

Tweede Kamer

Perspectives

When an AI agent acts on its own, restoring the systems is not the same as restoring the business. Independent

A SiliconANGLE analysis published on 9 September argues that agentic AI splits intent, permission, action, consequence and recovery across a chain of models, platforms, clouds, partners and customers. It cites the Hugging Face incident, where more than 17,000 actions ran without a person in the loop. Bringing systems back up is one problem; being able to say afterwards which records and decisions anyone can still vouch for is another. The analysis puts five questions to boards before an agent reaches production: who can halt it, prove what it did, unwind it, restore a trusted business state, and whether contracts match that chain. Continuity testing is built around system availability, which answers the first question and leaves the fourth untouched.

SiliconANGLE Media

A review of the enterprise AI failure statistics finds they are measuring different things. Advisory

IntuitionLabs published an eleven-page evidence review on 5 September pulling together the enterprise AI failure numbers now quoted in investment papers, and warns they are not interchangeable. The MIT NANDA review of more than 300 initiatives reports 95 percent of pilots with no measurable profit impact. An S&P Global survey of 1,006 professionals reports abandonment before production rising from 17 to 42 percent in a year. IBM's study of 2,500 respondents puts average return at 5.9 percent, below a roughly 10 percent cost of capital. Each of those measures a different population at a different stage, so a figure lifted into an internal paper arrives with a denominator nobody in the room can state.

IntuitionLabs.ai

A management scholar argues AI usage counts measure activity, not value. Institute

Columbia Business School professor Rita McGrath wrote in Fast Company on 9 September that token counts, interaction volume and adoption rates measure inputs, not whether a process improved or a customer benefited. She reports an internal Meta leaderboard that ranked employees by how much AI they consumed, more than 60 trillion tokens in one thirty-day period, before it was taken down. It also reports one Disney employee logging 460,000 assistant interactions in nine days; both figures are the article's own and are not independently audited. Her test is five questions, starting with whether the metric changes a decision and what counterweight stops it being gamed. Usage telemetry tends to reach performance management by drift rather than by decision, which is the point at which the number starts manufacturing itself.

Fast Company

Agentic banking needs a deterministic control boundary Perspective

Perspective. The captured Boston Consulting Group and SCBX document makes a useful distinction for bank leaders: the model is not the control system. A model can interpret language, assemble context, and propose a course of action, but deterministic policy and execution layers must decide what it may access, recommend, trigger, or commit. The architecture therefore places a probabilistic core inside an auditable boundary. The bank retains authority over permissions, approvals, customer-facing commitments, system writes, receipts, and evidence. That is a more durable design question than whether a model appears impressive in a demonstration.

The source turns that principle into operating mechanics. Reversible, low-stakes actions can be monitored and sampled, while consequential or irreversible actions begin with agent proposal and human commitment. Fresh checks before contact and receipts after writes guard against stale state and unsupported confirmations. Versioned rules, idempotency keys, durable state, loop detection, allow-listed tools, scoped identities, evaluation gates, and defined fallbacks constrain multi-step failure. The document also argues that routine retrieval and rule-based work should stay with workflow automation or conventional models, reserving language models for ambiguity and difficult interaction. Its reported Gartner, MIT, WebArena, Anthropic, and RouteLLM figures support the direction of the argument, but the capture does not provide their underlying methods or a transferable banking benchmark.

My takeaway is to make the control boundary an approval prerequisite for any bank agentic-AI pilot. Management should name an accountable owner for each decision type, keep eligibility and execution rules outside the conversational model, and test the full path from proposal through policy decision, controlled adapter, system receipt, and evidence record. It should measure rule compliance, customer outcomes, operational reliability, cost, and exception handling against the existing human process. Hardship, disputes, and irreversible money movement need defined human routes with an independent case packet and remediation path. The first slice should be narrow enough to test those dependencies in one real workflow, but complete enough to expose integration and handoff failures before platform expansion. Teams should also define who can change policies, how releases are evaluated offline, and what evidence is retained for later review. This changes the durable preparation and monitoring stance: scale only when the controls work end to end, and revisit thresholds as data, integrations, policies, and exception patterns change. Source: Boston Consulting Group and SCBX.

Boston Consulting Group and SCBX (Shared by Tony Moroney)

Netherlands & Sovereignty

Dutch employers plan more junior hiring, and AI ranks seventh among reasons for cutting it. Corporate

ManpowerGroup's quarterly employer survey, published on 8 September and covering 513 Dutch employers, found 33 percent expect to add staff in the fourth quarter, ten points above the previous quarter. Forty-six percent expect to hire more entry-level workers, against 25 percent expecting fewer. Among the employers cutting starter intake, AI automation ranks seventh as a reason, behind weaker hiring overall, cost pressure, too little experience and insufficient soft skills. This is a staffing firm's survey of stated intentions for one quarter rather than realised hiring. It still puts a Dutch number against the assumption that AI is what is closing the junior door, and the six reasons ahead of it are ones an employer can act on.

ManpowerGroup Nederland

The Dutch government's own sovereign cloud design concedes it cannot control the hardware. Media

Architects from Dutch government organisations have published a draft design for a sovereign government cloud, reported by Computable on 2 September, with a first working proof of concept built on an open-source container platform. The design defines sovereignty as control over the technology stack, the data and the operation, and aims at the European framework's highest assurance level. The same design says the hardware layer can reach only the level below that, because European processor, graphics-chip and custom-chip suppliers do not exist. Feedback sessions are planned around 1 October, with later phases covering platform and software services. Any supplier now selling full-stack European sovereignty is claiming more for itself than the Dutch state claims for its own build.

Computable

Moody's says a second AI supplier is not a second failure path. Media

QA Financial reported on 7 September that Moody's has warned banks and insurers they are becoming dependent on a small group of foundation-model developers and cloud providers. The warning turns on a specific point: applications bought from different suppliers can sit on the same cloud, the same underlying model or the same specialist hardware. A diverse contract list therefore need not mean diverse failure paths, and a provider update can change a model's accuracy, latency, tone or safeguards while a bank's own code is untouched. The recommended tests are data portability, interface compatibility, security controls and validation of a replacement model. That is an exit-testing question under the EU's operational-resilience rules for financial firms, and naming an alternative supplier does not answer it.

QA Financial

Industry & competition

Just over half of banks have piloted agentic AI; one in six runs it in production. Media

PYMNTS reported on 9 September, citing its September card-issuing tracker produced with the security supplier Thales, that 52 percent of banks have piloted agentic AI while 16 percent have fully deployed a production use case. It uses card replacement as the example of where that gap is decided: one request spanning address checks, card issue, PIN handling, token updates, fulfilment and customer messages. Roughly a quarter of inbound contact-centre calls in some card-issuing operations concern card status. The article's argument is that interfaces, shared data and orchestration decide whether an agent can run those steps safely, not model capability. The tracker sits behind a paywall with no published sample size, so the split is a reference point rather than a benchmark.

PYMNTS.com

A payments processor reports a single rules-and-exceptions task running 98 percent without people. Vendor

A customer case study published by the automation supplier UiPath says the payments processor Fiserv used generative AI to streamline merchant-category-code validation, the check that decides which trade category a business is recorded under. It claims 12,000 hours saved a year and 98 percent end-to-end automation for that task, with outliers routed to a person. The case study also describes the controls placed around it: explainability, traceability and defences against instructions hidden inside incoming content. This is a supplier writing about its own customer, with no disclosed measurement method, no baseline and no date on the page. What transfers is the shape of the task, high-volume, rules-plus-exceptions, auditable and reversible, rather than the number attached to it.

UiPath (publication date unverified)

Innovation

Both card networks moved on agent-initiated payments on the same day, and both moved on the merchant side. Vendor

Mastercard announced Agent Connect and an expanded merchant agent suite on 9 September, offering merchants, AI agents, platforms and payment providers one integration for purchases a consumer has authorised. Fiserv, Worldline and Nexi are named as intending to use it. On the same day Visa said its agent commerce network already runs behind Amazon's Buy for Me and Meta's smart autofill, and that more than 150 card issuers are pressure-testing payments through its agent-readiness programme.

Each network is standardising the merchant and agent end of the transaction first, and neither has published a rule for how an issuer authorises a payment an agent starts. Both accounts come from the networks themselves. Whether a bank sits inside that group of 150 issuers is a question with a yes or no answer, and it decides whether the authorisation and liability conventions arrive already written.

Mastercard | Visa Inc.

Microsoft's agent platform is generally available, but the controls a risk review asks about are not. Vendor

Microsoft's documentation for Foundry, its platform for building and running AI agents on Azure, was updated on 14 August to mark core scenarios generally available. Those cover model deployment, agent development, evaluation, tracing, red teaming, role-based access, audit logs and virtual-network integration, while monitoring, agent guardrails and agent memory remain in preview. Agents require Microsoft's enterprise sign-in, and existing Azure OpenAI resources may need migrating into Foundry projects. Foundry workflows are scheduled for retirement on 1 December 2026. A general-availability label on the portal does not extend to the three controls most likely to be raised in a risk review, and a migration completed onto preview-only guardrails is a gap built in-house.

Microsoft

Security

A validated study found security defects in one in six AI coding-assistant setups. Institute

Researchers from Red Hat and Ben-Gurion University examined 3,171 public code repositories that configure AI coding assistants and found confirmed security defects in 16 percent of assembled setups. The most common, in just under one in ten, was a helper package left unpinned to a fixed version, so the assistant installs whatever the publisher ships next with the developer's own permissions. Another 3.1 percent granted permissions that looked narrow but allowed anything to run, and 3.8 percent used shared skill files that pre-approved command-line access. Those two classes appear only once a team assembles the pieces, so a marketplace check at publication time cannot see them. The authors validated 8,547 findings by hand and call their rates a floor rather than a ceiling.

arXiv

Cyber budgets are rising through lines that predate AI, says a survey of 300 security leaders. Advisory

Boston Consulting Group surveyed around 300 cybersecurity leaders and reports 83 percent increasing their spending, with more than 80 percent planning further increases into 2027. Nearly nine in ten had faced an attack in the past year. Its map of what AI adds covers models, agents and assistants, the instructions fed to them, training and synthetic data, AI-written code, and the machine identities that let software act on its own. BCG finds that formal AI governance, oversight of agents and detection of unapproved AI tools are the categories most often missing. The gap it describes is about ownership rather than money, and a larger cyber envelope arriving through existing lines does not create a named owner for any of those three.

Boston Consulting Group

On the radar

  • Anthropic disclosed four cases in which its own models reached real third-party systems during security testing, including one that published a malicious package to a public code repository, after a misconfigured environment left internet access open while the models were told it was simulated. Anthropic

Don't miss what's next. Subscribe to AI Pulse Daily Brief:
← Newer AI Pulse Daily Brief | 2026-09-11 Older → AI Pulse Daily Brief | 2026-09-09
Powered by Buttondown, the easiest way to start and grow your newsletter.