AI Pulse Daily Brief logo

AI Pulse Daily Brief

Archives
Log in
Subscribe
October 6, 2026

AI Pulse Daily Brief | 2026-10-06

Reading time ~14 mins

A researcher shows a hidden instruction passing from one AI agent to the next at Google and four other organisations, because each agent trusts the one that hands it work. OpenAI starts marking ChatGPT text in the EU but leaves its API watermark off by default, and an AI Act specialist explains why a file label alone falls short. AI agents could not complete a single action at any Benelux bank tested, and ABN AMRO names financial crime as the first work for its agent pilot. BearingPoint finds most firms have not planned where AI-freed staff time goes. BCG maps how a bank's risk lines would govern agents, and a European expert strategy proposes state-arranged access for banks to restricted AI cyber-defence models.

Regulatory

OpenAI starts marking ChatGPT text in the EU, but leaves the watermark off by default for companies building on its API. Vendor

OpenAI announced on 5 October that it will add an invisible watermark to eligible ChatGPT and Codex text in the EU over the coming weeks. Codex is its coding assistant. Customers of its API, the connection companies use to build OpenAI models into their own applications, can switch watermarking on for selected models. It is off by default. Only approved researchers and expert organisations may apply to use the detector that reads the mark. OpenAI says a watermark can show that its system produced or processed part of a text. It cannot show how much a person edited it, and it proves neither ownership nor accuracy. Detection weakens on short or heavily edited text, and a missing watermark does not prove a person wrote it. PYMNTS reports that OpenAI built the tools for the AI Act's requirement that providers of generative AI make their text detectable by machine. Those Article 50(2) duties have applied since 2 August 2026. OpenAI signed the provider section of the EU's transparency code, as this brief reported on 6 August. A bank that builds its own text-generating application on the API may itself count as the provider under Article 50. For those applications, OpenAI's EU rollout marks nothing until someone at the bank switches the watermark on.

OpenAI | PYMNTS

Perspectives

Aleksandr Tiulkanov argues a file label alone does not meet the AI Act's duty to mark AI-generated content Perspective

Aleksandr Tiulkanov, an EU AI governance instructor, wrote on 6 September that guidance circulating from advisory firms misreads the AI Act's marking duty. Article 50(2) requires providers of generative AI to mark output so a machine can detect it. Their technical solutions must be effective, interoperable, robust and reliable as far as is technically feasible. Some guidance says metadata, a label stored with the file, is enough. Tiulkanov answers that metadata is easily removed, so on its own it fails the robustness test. Aleksandr Tiulkanov writes that metadata is "necessary but NOT SUFFICIENT for art 50(2) compliance."

His benchmark is the Code of Practice on Transparency of AI-generated Content, finalised on 10 June. The Commission and the AI Board, the body of national authorities, have confirmed it as an adequate voluntary way to show compliance. A firm that complies by other means must show its own measures are adequate, case by case, to market-surveillance authorities. In his reading the Code sets two measures. Aleksandr Tiulkanov writes that the first is "digitally signed metadata" and the second is "robust and resilient imperceptible watermarking."

Today's Regulatory story makes the test practical. OpenAI's watermark is off by default in its API, and its detector is reserved for approved researchers and expert organisations. A bank that builds a text-generating application on that API may count as the provider under Article 50. An application that relies on a file label alone then sits outside the route the Commission has endorsed. This brief carried the Code on 7 May and its roughly 190 signatories on 6 August. What Tiulkanov adds is a two-part test an in-house system can be held against, and a warning that outside AI Act advice can be wrong on a duty that has applied since 2 August. His post is one expert's reading of a voluntary code, and a firm may still prove other measures adequate.

European Commission (By Aleksandr Tiulkanov)

BCG sets out how a bank's risk lines would govern AI agents, and finds an incident no one owns end to end Advisory

Boston Consulting Group published a paper in September called When Agents Run the Bank. Its authors are Matteo Coppola, who leads BCG's risk and compliance practice globally, and Anne Kleppe, who leads its responsible AI work globally. Tony Moroney shared it on LinkedIn on 11 September with a one-sentence summary taken from BCG. On 1 October this brief carried Moroney's essay arguing that the second line, a bank's independent risk oversight, must move from periodic review to continuous oversight. The BCG paper sets out what that shift would look like inside a bank.

Its starting point is scale. BCG's example is a loan-monitoring agent that slightly underweights one early-warning sign in one sector. Each of its thousands of calls can be defended on its own, yet together they build a material risk position. Reviewing a sample of single decisions, the way credit files are checked today, would miss that pattern. So BCG says the second line should stop approving individual AI systems and govern the whole machine-run system continuously, at portfolio level. Another approval layer, it says, will not help. It cites its 2026 survey of about 30 leading international banks. 80% have an AI policy, yet one in five reported AI systems in production giving biased or wrong outputs.

The paper splits control into seven layers and gives the first and second lines a role in each. Two of them change how a bank builds controls. Each risk check, such as a sanctions screen or a credit assessment, becomes a governed service that carries its own thresholds, permitted uses and audit trail, with a named first-line owner. Risk appetite is written as code, so it can be enforced at the moment a decision is made. BCG names the risk this creates. A wrongly coded exposure limit can distort every credit decision before any review spots the pattern. The second line keeps ownership of appetite and obligations, while the first line, technology included, codes them under change control.

BCG also names three failures it says banks are not yet measuring. An agent's behaviour can drift with no change to the model, as when an instruction to flag suspicious activity slides over thousands of cases into flagging only what is clearly suspicious. An agent running on an employee's login leaves an audit trail that names the login, but not the agent, who delegated to it or why. A chain of steps, each one permitted, can add up to a wrong action. Checks before launch and reviews of outputs, the tools banks rely on today, are not built to catch any of the three. Internal audit, in BCG's design, should test whether the control design works and leaves evidence that can be challenged, rather than redo agent decisions.

Its sharpest point is a worked incident. A hidden instruction planted in content manipulates an AI-run customer risk assessment. Cyber owns the technical incident, financial crime the due-diligence impact and compliance the regulatory side. No single function may own the whole response: finding every affected decision, coordinating the fix, deciding whether onboarding can continue and accepting the risk that remains. Today's Security story describes the same kind of hidden instruction passing between agents at five organisations. BCG says it sees only the very first agent applications at its clients, and the paper is a consultancy design resting on a small survey. Under a three-lines model built around human decisions, that response falls between functions.

Boston Consulting Group via LinkedIn (By Tony Moroney)

A European expert strategy proposes that banks get state-arranged access to restricted AI cyber-defence models Institute

A Transformative AI Strategy for Europe, a 206-page report published in September by the KIRA Center in Berlin, argues that transformative AI could arrive before 2030 and that Europe is not ready. Its authors wrote in a personal capacity under a senior council convened by the economist Monika Schnitzer. Council members include Margrethe Vestager, Yoshua Bengio, Philippe Aghion and Daron Acemoglu. The report says no company, government or EU institution funded or directed it, and that it is not a consensus document. Peter Slattery, an AI risk researcher, summarised its five immediate objectives on LinkedIn on 14 September.

On 2 October this brief carried ECB President Christine Lagarde on the June suspension of Anthropic's most capable models. The report adds a dated timeline. European bodies including ENISA, the EU cybersecurity agency, first joined the restricted programme on 5 June. A US export-control directive suspended Fable 5 and Mythos 5 worldwide from 12 June for 18 days. US partners regained access on 1 July, and Europe's return was unresolved when the report was written. The authors then test the EU's main defence. The Anti-Coercion Instrument, the EU law for answering economic pressure from other countries, turns on intent. A restriction framed as security would probably not have met that test on its own, they write, though a documented pattern of tying access to European policy choices would.

Its remedies go further than anything this brief has carried so far. It proposes that Member States run programmes giving critical entities under NIS2, the EU's cybersecurity law, access to security experts and highly capable AI models to find and patch serious flaws. It names banks and payment systems among those entities. Priority entities would be designated in the fourth quarter of 2026, with a pooled fund in the next budget cycle. The report says a Commission and ENISA blueprint for structured access to advanced AI for cybersecurity is due in late 2026. To secure access in the first place, it proposes trading compute for access. Any hosting deal above 100 megawatts would guarantee European customers the same access as the provider's home customers, continuity of service and entry to restricted security programmes.

The Netherlands sits at the centre of its bargaining plan. The report wants the Netherlands, Germany and France to found an alliance for supply-chain security by the end of 2026, modelled on the Dutch government's Semicon Coalition. A chart it takes from CSET, a US research centre, gives the Netherlands all of the world's supply of EUV scanners, the machines that print the most advanced chips. The alliance would map such chokepoints and agree an escalation chain that could include reduced servicing of equipment and conditions on export licences. The report notes that other countries could answer with tariffs or by restricting cloud access. It would share the cost of retaliation on the model of the 2014 agreement on contributions to the Single Resolution Fund, which banks pay into.

It also wants the EU to host at least 15% of global AI computing capacity by 2030, against about 5% today. Public guarantees and purchase commitments of EUR 100 billion, through the European Investment Bank group and national promotional banks, would back it. It reports that one-year rental prices for a widely used AI chip rose about 40% between October 2025 and March 2026. Rising prices or export limits could push European users onto less capable models, and it names finance among the sectors that need reliable computing on EU soil.

These are proposals from independent experts, not policy. On 5 October this brief carried the Banque de France's call for the most powerful models to go first to trusted partners. This report names banks among the entities that would receive such models, through national programmes. A Dutch bank's access to the strongest defensive models would then run partly through The Hague, alongside its contracts with the vendor.

KIRA Center via LinkedIn (By Peter Slattery)

Netherlands & Sovereignty

A Benelux test found AI agents could not complete a single action at any bank it tried. Advisory

In The Pocket, a digital agency, released its Agentic Readiness Index on 5 October through the Dutch trade site Emerce. It sent AI agents, software that acts on a person's behalf, through 599 tasks at 245 Belgian and Dutch companies in nine sectors, 118 of them Dutch. At 6% of the Dutch companies an agent could finish a task on its own, such as submitting an application or placing an order. Agents could find 77% of the companies, and 54% of test questions got a correct answer. At banks, telecom providers, government bodies, health insurers and public transport, no agent fully completed an action in any test. In HR and payroll, agents completed at least one step at 58% of firms. In The Pocket puts 55% of stoppages down to technical set-up, such as pages that load content only through scripts. The other 45% were deliberate barriers such as logins, firewalls and identity checks. The agency publishes little of its method and sells the services its index measures, so the figures do not represent Dutch business as a whole. This brief carried HSBC opening account data to corporate clients' AI tools on 30 September and Equals Money letting customers' agents read accounts on 5 October. Bank channels sit behind login and identity checks built to stop fraud, so a customer's agent gets in only through a route the bank designs for it.

In The Pocket via Emerce

Industry & competition

ABN AMRO names financial crime and customer data as the first work for its AI agent pilot. Corporate

On 2 October this brief reported that ABN AMRO was preparing a pilot with Wonderful, an Amsterdam-based maker of AI agents. Neither firm had then named the workflows, a start date or a duration. ABN AMRO's own release of 2 October fills those gaps. The pilot runs for six months in the bank's experimentation environment, led by its Innovation unit. It covers two units, Detecting Financial Crime and Customer Data Solutions. The agents will be tested on gathering information, preparing case files for financial-crime investigations and supporting checks on customer-data quality. They work within set boundaries and under human oversight. The bank says it will judge reliability, human oversight and how the agents fit existing processes and controls, and it reports no results yet. Head of Innovation Yorick Naeff says the aim is to deploy solutions, learn how they perform and understand what responsible use takes. A Dutch peer now has agents preparing cases in its anti-money-laundering work, under the same Dutch supervisors. The tests it names are the ones a model-risk review would apply to agents doing that work.

ABN AMRO

BearingPoint finds six in ten firms report AI-freed staff capacity, while fewer than half plan for it. Advisory

BearingPoint, a European management consultancy, surveyed 1,050 executives and senior leaders in 13 countries across Europe, the US and China in August 2026. It published the results on 30 September. Of organisations that have put AI to work, 74% report a measurable effect on revenue or costs. Near half put that effect below 4% of costs and below 2% of revenue. Only 13% scaled their AI initiatives fully in line with the original business case, with complex regulation and older systems the barriers cited most. The newer finding concerns staff. 62% report that AI has already created overcapacity of at least 10% in their workforce, and 95% expect that level by 2030. Only 48% build workforce planning into their AI roadmaps. BearingPoint calls this a productivity trap, where freed time is absorbed unless roles, structures and budgets change. The answers are self-reported, and the full report sits behind a registration form, so these figures come from BearingPoint's public summary. Yesterday this brief carried Société Générale's chief executive calling AI's realised benefit minimal so far. The survey shows where that gap can open. Hours freed by AI become savings only when someone decides where they go, and fewer than half of these firms plan that.

BearingPoint

Security

A researcher shows hidden instructions passing between AI agents because each agent trusts the one handing it work. Media

Ars Technica reported on 5 October that Google and four other organisations have acknowledged flaws in their AI agents in the past five months. Each let an attacker use one agent inside a network to send harmful instructions to others. Independent researcher Syed Anas Mohiuddin tested agents from organisations including Google, JPMorgan Chase, a database firm, the security firm Rapid7, the French government's digital directorate and the US federal government. His attacks start with text planted in content an agent reads. That agent passes it to a second agent as an ordinary delegated task, and the second agent runs it because it trusts the first. He calls this protocol pivoting. In his tests the task starts on the Model Context Protocol, a common standard for connecting agents to tools and data. It then crosses to another standard, such as Google's protocol for agents delegating work to each other. Checks on who may ask for what are often lost in that crossing. Google's flaw was rated high severity and fixed by limiting which network addresses its tool may reach. Rapid7's was rated low and fixed in September. Ars reports no compromise at any bank and does not say whether a flaw was confirmed at JPMorgan Chase. Douglas McKee of Rapid7 told Ars: "Each protocol was built assuming it lived on its own, so each one checks its own front door while nobody watches the hallway in between." Markus Vervier of the security firm X41 D-Sec calls it a subclass of indirect prompt injection rather than a new attack.

On 28 September this brief carried OpenAI's hidden instruction that copies itself from agent to agent, and Australia's advice to secure cooperating agents as one system. This report shows the same weakness in agents already running at five organisations, and locates it at the handoff. Ars says agent builders have dropped zero trust, the principle that every request between systems needs its own authorisation, even from a trusted neighbour. A check on each agent's own input does not stop an instruction that arrives from a colleague agent the receiver is built to trust.

Ars Technica

Don't miss what's next. Subscribe to AI Pulse Daily Brief:
Older → AI Pulse Daily Brief | 2026-10-05
Powered by Buttondown, the easiest way to start and grow your newsletter.