AI Pulse Daily Brief | 2026-09-03
Reading time ~11 mins
The Financial Stability Board tells G20 ministers that frontier AI’s effect on cyber risk is the
financial system’s most immediate concern, days after 90 days of decoy systems showed AI gateways
being probed for stored credentials. France and the Netherlands ask Brussels for European-content
rules in public buying. TNO gives the Dutch government a working method for judging whether a
language model is genuinely open and European. Only 35% of finance leaders say they can measure
what AI returns.
Perspectives
Gary Marcus warns a reported OpenAI change would hide most of a model’s reasoning. Skeptic
Writing on 2 September, Gary Marcus responded to reporting by The Information that OpenAI is exploring a technique that reveals less of a model’s chain of thought. That is the step-by-step working a model exposes as it answers. He argues that watching that working, imperfect as it is, remains one of the few practical ways to see what an opaque model is doing. His recommendation is to keep the visibility until the performance gain and the safety cost have both been independently measured. A change of this kind reaches a bank through a routine vendor version upgrade, and nothing in a standard model-change record would prompt anyone to notice that a monitoring control had quietly gone.
Ed Zitron assembles the disclosures showing AI infrastructure revenue rests on a few slow payers. Skeptic
In a 1 September essay, Ed Zitron argues that the AI building boom is being read as proof of durable demand before anyone has shown diversified, profitable cash flows behind it. He cites NVIDIA disclosures that five customers accounted for about 70% of the money owed to the company, and three for about 44%. Over the same quarter, the average time to collect payment stretched from roughly 45 to 60 days. He sets that beside CoreWeave, which grew revenue about 112% in a year while posting a quarterly loss near $646 million against roughly $24 billion of debt. These are ordinary counterparty-concentration and collection signals sitting inside technology purchasing decisions, where the credit function is not usually asked to look.
AI productivity is an operating-model redesign problem Perspective
Perspective. McKinsey's August 2026 article makes a useful distinction for bank leaders: meaningful AI impact in product and software development does not come from adding copilots to unchanged routines. It comes from redesigning the full product-development system. The underlying survey of 334 product and engineering leaders reports uneven outcomes: only 25 percent of director-and-above respondents reported meaningful or top acceleration, while 30 percent reported falling productivity. The organizations pulling ahead were described as changing workflows, roles, verification, AI operations, and change management together.
The evidence points to a concrete mechanism. Faster generation can amplify ambiguity, handoffs, technical debt, and unsafe changes unless product intent is translated into clear requirements, acceptance criteria, permissions, telemetry, review gates, and accountable decision rights. The article reports that 93 percent of top accelerators embedded AI into workflows, and that organizations redesigning processes before introducing technology were more than twice as likely to report productivity gains above 20 percent as those layering AI onto existing ways of working. It also describes a Sonar case in which pull-request cycle duration fell 3.4 times, throughput improved 2.2 times, and build activities became 50–80 percent more productive; these are reported examples, not universal benchmarks.
My takeaway for a bank is that the next AI investment decision should be framed around work-system readiness, not licence volume. Start with a small number of high-value workflows and map triggers, handoffs, human judgment, controls, telemetry, and customer outcomes end to end. Then measure quality, security, reliability, cost, time to market, rework, and review burden alongside productivity. A shared AI-operations layer should make identity, permissions, evaluation, logging, escalation, policy enforcement, and cost visible, with stronger evidence and human approval for regulated or customer-facing changes. The operating model should also clarify which decisions agents may recommend, draft, or execute, which evidence a reviewer must inspect, and how exceptions move to a named accountable owner. Teams need coaching and protected time to practise, while leaders should review outcome measures rather than prompt counts or code volume. That creates a preparation stance for procurement and deployment: fund controls, measurement, role redesign, and change management as part of the capability, not as remediation after rollout. This is a durable monitoring stance because the article's survey and cases do not prove causation or guarantee transfer to banking; they do show why tool adoption without operating-model redesign can leave faster work producing faster risk.
McKinsey & Company (Shared by Tony Moroney)
Netherlands
TNO hands the Dutch government a four-part method for judging open and European language models. Institute
TNO Vector published an English summary on 1 September of a report written for the Ministry of the Interior and Kingdom Relations on choosing open and European language models. It treats openness as a spectrum of five levels, running from closed systems through open weights to fully open, rather than as a label a supplier can simply claim. European provenance is judged separately on three counts: who governs the provider, where data and infrastructure sit, and whether training covered European languages. A decision table combines those with technical fit and intended use, and a demonstrator ranks 19 public models against weighted priorities. This is now the working vocabulary of the ministry that writes Dutch digital-sovereignty policy.
Dutch workers are using generative AI well ahead of the support their employers give them. Institute
Researchers at De Nederlandsche Bank, writing in the economics journal ESB on 21 August, report that only 25% of Dutch workers experienced active employer support for generative AI in an October 2025 survey. In most sectors, worker use ran more than 15 percentage points ahead of active encouragement or training, and the gap approached 30 points in education and government. Three in ten workers using AI held a paid licence, and those were concentrated among the supported. Read that way, the licence figure is an estimate of how much workplace AI use is running on unlicensed consumer tools, which is where the data-handling exposure sits rather than in the training numbers.
Sovereignty
France and the Netherlands ask Brussels for European-content rules in public technology buying. Authority
The Dutch government published a joint statement with France on 2 September calling on Europe to cut its strategic dependencies across critical technology. It treats computing capacity as the foundation of AI sovereignty and ties it to energy, advanced semiconductors, photonics and quantum work rather than treating any of them separately. The two governments back the European Commission’s technology sovereignty package, targeted European-content criteria in public procurement and state aid, and a cloud framework shielding sensitive data from foreign legal reach. Supplier ecosystems reorganise around a public-procurement definition long before any obligation reaches a private bank, so the definition of European content is worth tracking while it is still being written.
The EU’s proposed chip shortage regime names banking among the sectors that can trigger it. Advisory
The law firm Freshfields published an analysis on 27 August of the European Commission’s June proposal to replace the 2023 Chips Act. The proposed crisis regime runs in three stages, from mapping and monitoring to alert and activation and then to shortage response, and it lists banking, financial-market infrastructure and digital infrastructure among the critical sectors. Freshfields notes that the EU makes under 10% of the world’s semiconductors and is almost wholly dependent on the United States and Asia for the most advanced ones. Activation would require concrete evidence that a critical sector is affected, so an institution named as critical but unable to evidence its own chip dependency holds a claim it cannot exercise. A crisis blueprint is planned for the second quarter of 2027.
A 165 megawatt AI data centre in Finland arrives tied to one American chip supplier for seven years. Vendor
Cerebras Systems and Compute Nordic announced on 1 September a planned AI data centre in Mikkeli, Finland, with 165 megawatts of contracted capacity. Construction has started on a 50 megawatt first phase, with later stages at 80 and 165 megawatts, under seven-year service orders and an estimated regional investment of €1.0 to €1.7 billion. The capacity is European; the accelerator technology, software and support behind it are not, and the release says nothing about access to other suppliers’ chips or an exit path. Location and control are separate tests, and a long single-supplier service commitment reproduces inside the EU the concentration that European capacity is usually bought to reduce.
Industry & competition
Only 35% of finance leaders say they can confidently measure what AI returns. Media
CFO Dive reported on 27 August on Protiviti’s 2026 Global Finance Trends Survey of 902 finance executives worldwide. It found 77% of finance organisations using AI, 35% confident they can gauge the return, and 14% working to a defined AI strategy. Protiviti’s own summary says organisations are adopting AI faster than they can measure it, and that automation, analytics and modernisation are delivering stronger measurable value than AI alone. The sharpest finding is not that AI value resists measurement but that it currently loses the comparison against those alternatives, and that lost comparison is the objection an investment case meets at the approval gate.
A German private bank is moving from one custom assistant to bank-wide AI agents in stages. Vendor
Google Cloud published a customer account describing how Berenberg, a German private bank, built a custom assistant called BegoChat and is now moving to a phased platform rollout starting in 2026. The design keeps both general and role-specific agents centrally managed, with training, monitoring, human review and no fully automated decision-making. Google reports morning-briefing preparation running 85 to 90% faster and about an hour a day returned to client conversations, figures the vendor publishes with no baseline and no sample. The auditable content is the sequence and the stated no-automated-decision boundary, which is what a supervisor would ask about; the time savings carry no evidence at all.
Google Cloud (publication date unverified)
Innovation
Microsoft says the biggest cost in running AI agents is finding the right tool, not the model. Vendor
Microsoft published guidance on 2 September for its enterprise AI platform arguing that what an agent is fed on each turn drives both cost and quality. It describes retrieval that runs several sub-questions at once and returns citations, plus centrally managed memory of a session, a user and a procedure. Agents reach their tools through a single managed endpoint that carries authentication and access policy. Microsoft’s own evaluations report up to 54% better evidence recall, 34% lower retrieval cost, and around 97% less input consumed when agents search a large tool library rather than loading all of it. These are vendor figures, but they point the cost question at the size of a pilot’s tool catalogue, which no choice of model would fix.
Snowflake is selling a managed control plane for agent access, turning a build question into build-or-buy. Vendor
Snowflake published an enterprise guide on 2 September positioning its Cortex AI Gateway as a production control plane for how agents reach systems and data. The design hands out credentials at run time so agents never hold production secrets, binds access to a named user’s identity, applies policy to individual tools, and keeps an audit trail. Its recommended first step is an inventory of the unmanaged tool-connection servers already running inside the organisation, followed by routing everything through one gateway with approved-server lists, rate limits and data-loss controls. This is vendor positioning with no named bank deployment, but the inventory step is the only part whose value does not depend on which control plane is eventually chosen.
Security
A global financial standard-setter and a cloud security body land on the same warning two days apart. Authority
The Financial Stability Board’s chair told G20 finance ministers on 31 August that frontier AI’s effect on cyber risk is the financial system’s most immediate concern. Increasingly autonomous models, the letter argues, change the speed, scale and economics of that risk. Two days later the Cloud Security Alliance turned 90 days of observed activity against AI infrastructure into a control checklist. Its list runs to inventorying the AI stack, keeping test endpoints off public networks, binding access to identity, restricting outbound traffic, rotating exposed keys and patching on a production cadence.
One body warns about what frontier models could do; the other records what is already being tried against the software running underneath them. A standard-setter’s letter to G20 ministers is how supervisory questions reach national workplans, so this lands inside the next operational-resilience scenario review rather than as a separate AI exercise. The Alliance also states its note was written with AI assistance and never passed the Alliance’s own review, which makes it a checklist to test rather than evidence to cite.
Financial Stability Board | Cloud Security Alliance
A study traces an AI model failure through to modelled bank failures in a payment system. Institute
A paper posted to arXiv on 24 August works through the United Kingdom’s real-time payment settlement system to show where model-level testing stops short of explaining harm to the wider system. Structured hazard analysis produced 104 unsafe control actions unrelated to AI, plus eight loss scenarios that AI could drive. Simple adversarial text then measurably shifted a stylised AI-assisted portfolio allocator, and the authors mapped that shift into a 50-bank contagion model, where it raised modelled forced-selling intensity and failure rates under stress. They are explicit that this is illustrative rather than a forecast. The transferable part is the ordering rule: a benchmark score is not a safety case until the path to a named business loss has been written down.
arXiv: From model failure to bank-system harm
On the radar
- Security researchers at Wiz ran decoy AI gateways, agent frameworks and vector databases for 90 days and watched intruders move from an exposed tool endpoint through to the gateway’s stored provider credentials. Wiz Threat Research