AI Pulse Daily Brief | 2026-09-16
Reading time ~12 mins
A critical flaw in a widely used AI tool connector shows that a patched server can still work as a route into internal systems. EY's first AI governance survey finds a quarter of organisations unable to detect AI agents they never authorised. Two independent research groups measured what autonomous agents actually produced and found almost nothing. Prinsjesdag puts €120 million into a European AI project while the new Dutch innovation agency starts on €10 million. Dutch data-centre permits now take four and a half years, and grid connections up to ten.
Top signal
A patched AI tool connector can still work as a route into internal systems. Institute
The Cloud Security Alliance published a nine-page research note on 8 September about two chained flaws in an open-source connector that lets AI assistants query Grafana, a widely used monitoring dashboard. Before the fix, the connector accepted a session identifier as proof of who was calling, and one of its tools let that caller choose the destination, method and body of an outbound request. An unauthenticated caller could therefore use the connector's network position to reach internal services and the cloud service that hands out credentials. Grafana Labs shipped a fixed release with optional token authentication eight days after the reports arrived.
The patch closes the inbound authentication gap and leaves the outbound path a configuration choice, so a server that is fully up to date with an open egress path is still usable as a proxy. The note cites an April 2026 scan that found more than 12,500 of these connectors reachable from the internet, and says it was itself produced with AI assistance without passing the alliance's own review. What limits the damage on each connector the bank runs or pilots is the egress rule attached to it, not the version number.
Perspectives
An accounting test on the AI building boom puts the break-even bar at a 2.7-fold productivity gain. Media
MIT Technology Review published a calculation on 15 September by the finance professor Jessica Wachter. Data-centre spending by the large cloud providers approaching $1.1 trillion by 2027 would need a 2.7-fold productivity improvement at the AI companies to break even by 2030 under her assumptions. The article sets three linked conditions behind that figure, which are provider revenue, broad economic productivity, and continued demand for expensive frontier models against cheaper ones that may be good enough. It notes that data-centre borrowing now runs through lenders, guarantors, private-credit funds, pension funds and insurers. It also cites a survey in which about 90 percent of senior executives reported no productivity increase over the previous three years.
Anthropic's chief executive offers outside evaluators permanent inside access with the right to publish. CxO voice
Fortune reported on 12 September that Dario Amodei's essay calls for a slower pace of frontier-model capability improvement. Its concrete proposal is permanent employee-level access for independent evaluators, with authority to verify safety practices, report incidents and publish findings without the company's editorial control. Fortune says Anthropic will implement the evaluator-access step immediately. The essay also asks for common standards among democratic countries and government coordination with authoritarian states on shared constraints, using AI-assisted biological weapons as its example. This is the first frontier-provider safety claim a buyer could actually check, because the absence of any published evaluator finding a year from now would itself be evidence.
Three security-literate authors take apart the claim that AI agents could take over the internet in six months. Skeptic
Gary Marcus, Nathan Hamiel and Zack Korman published a joint piece on 12 September arguing that an internet-wide takeover is vague and operationally implausible, because infrastructure, motive, cost and coordination are all left unspecified. They keep a narrower warning, which is that many individual systems are weakly defended and that models running on local hardware could lower the cost of attacking them. Their proposed release condition is that any agent which cannot be closely monitored should be restricted, with bounded permissions, complete action traces and a stop path that has actually been exercised. That is a testable version of a claim which is otherwise quoted at board tables in its cinematic form.
An executive AI assistant that removes disagreement makes the leader's existing view more persuasive. Institute
Two Northwestern professors, Julio Ottino and Brian Uzzi, argued in Fortune on 15 September that AI compresses two different kinds of friction. Coordination friction, made of scheduling, approvals and reconciliation, is waste, while cognitive friction, made of disagreement and inconvenient questions, is where challenge comes from. Their case is that an assistant which only makes information coherent can turn a leader's own assumptions into a more convincing answer instead of exposing the contradictions in them. They point to Meta stepping back from its most aggressive workforce-reduction plans after internal data questioned whether AI-assisted coding reached user-facing gains. The distinction leaves one question to ask of any assistant business case, which is whether the speed gain removes administrative drag or removes challenge.
A transparency lawsuit produced 132 pages on a US AI evaluation framework with the criteria redacted. Skeptic
Gary Marcus reported on 15 September that a freedom-of-information request and lawsuit by the group Protect Democracy produced 132 pages of United States records about a framework used to assess frontier AI models before release. He says most of the substantive criteria in that production remain blacked out, and quotes the group's counsel warning that secrecy raises the risk of both corruption and avoidable mistakes. Litigation to obtain the criteria and the justifications behind them continues. A safety label whose criteria cannot be inspected even under a court-ordered release is a gap in the bank's own assurance file rather than a control in it.
Netherlands & Sovereignty
Prinsjesdag put €120 million behind Dutch participation in a European AI project. Authority
The Dutch government's 15 September budget package commits €120 million to Dutch participation in an Important Project of Common European Interest on AI, the vehicle the European Union uses to co-fund cross-border industrial capability. The same package starts the National Agency for Disruptive Innovation and gives the state €500 million to act as a launching customer for innovative smaller firms. It also reserves €3.3 billion for a National Investment Institution and €18 million to study a Europe-Asia cable branch to the Netherlands. Procurement rules and implementation detail are not in the announcement. Which suppliers win those first reference contracts in 2027 is the cheapest early read on which European alternatives will be credible when sovereign-hosting requirements reach the bank's own procurement.
The Dutch justice minister says AI is cutting patching windows from weeks to hours. Authority
At Cybersec Netherlands on 10 September, Justice and Security Minister David van Weel warned that the country could slide unnoticed into what he called a digital 9/11 as attacks increase. He described AI compressing vulnerability and patching cycles from weeks to days and from days to hours, with offensive and defensive systems racing each other. He also noted that most critical services are run by companies rather than by government, and that defensive AI cannot make the human decision to install an update. If that framing becomes the national expectation, the gap a bank would be measured on is approval latency rather than detection.
The new Dutch innovation agency starts on €10 million against pilots costed at €30 to €50 million each. Media
Computable reported on 15 September that the National Agency for Disruptive Innovation is budgeted at €10 million in 2027, €25 million in 2028 and €75 million a year from 2029. Its own initiators had asked for at least €300 million in start-up capital and €150 million a year for five years. Three pilots planned for 2027 are each estimated at €30 to €50 million against that €10 million first-year budget, so they will be phased or only partly funded. The agency's programme areas are digitalisation and AI, security and resilience, and life sciences. Any Dutch AI consortium that cites this agency's co-funding over the next two years is relying on money the budget does not yet contain.
Dutch data-centre growth is now limited by the grid, permits and local consent. Media
Computable published an analysis on 11 September reporting more than 1.5 gigawatts of Dutch data-centre computing capacity at the end of 2025, against an industry baseline scenario of 2.8 gigawatts by 2031. It puts average permitting time at 4.5 years and grid-connection waits at up to ten years, and describes a 78 megawatt campus in Amsterdam-Westpoort leased by Microsoft. A proposal to place smaller compute units on existing infrastructure and at renewable-generation sites was accepted by a parliamentary committee as input for a debate in January 2027. At those lead times Dutch capacity is committed years before any procurement decision. That makes bringing workloads back a 2031 option rather than a contingency available if an EU sovereignty requirement lands sooner.
Industry & competition
Ryanair resolves four in five customer chats without a human, and published how it qualified the models. Media
PYMNTS reported on 15 September that Ryanair's AI handles 120,000 customer chats a day in seven languages and resolves 80 percent of them without a human agent. Amazon Web Services says the system has answered 10 million chats at 94 percent accuracy and cut response time from 18 seconds to 2.9 seconds. Combined with phone automation, the airline reports saving more than 500,000 agent hours and reducing agents per passenger by 70 percent. The transferable part is the qualification method rather than the rate, because Ryanair tested five foundation models against 12,000 real customer questions and keeps automated scoring on every live answer. In a regulated channel that scoring is the control, so a containment target quoted without it is an unevidenced claim.
Innovation
OpenAI launched a financial-services product aimed first at investment banking. Vendor
OpenAI introduced ChatGPT for Financial Services, aimed initially at investment banking and equity research, after design partnerships with Morgan Stanley and Evercore. The product combines financial data sets and a customer's existing data subscriptions with retrieval and financial reasoning, and produces valuation models, research notes and pitchbooks. PYMNTS reported the launch on 10 September as the first step before a wider financial-services rollout. The durable part is the route to market rather than the product, because anchor-bank design partnerships are how a vendor buys workflow specificity. A bank that later licenses the result inherits defaults that were settled inside another institution.
Google's newest real-time voice models complete about a third of banking tasks on its own benchmark. Vendor
Google announced Gemini 3.8 Live and an extended-reasoning variant on 15 September, both built for near-real-time spoken interaction with visual context and background tool calls. They are available now through Google's developer interfaces, with a private preview inside its enterprise product. Google reports 35.1 percent on Sierra's voice-banking benchmark, a published test of whether a voice agent completes a banking task from start to finish. That number is the one to carry into any voice-agent discussion, because at a third of tasks completed a contained internal-channel pilot and a customer-facing containment target are not the same decision.
Research
Two independent research groups measured what autonomous agents actually produced, and found almost nothing. Institute
dfdx labs ran a three-week experiment in which an agent operated a small business selling services to other agents. It built 17 paid products, three web pages and 46 programming endpoints, consumed about 5.4 billion tokens worth roughly $6,900 at list prices, and earned $1.54. A separate six-month study on arXiv followed autonomous trading across two live crypto fleets, about 7.5 million model calls and 300,000 recorded actions, and found no directional edge after fees. In that study, configuration such as risk-slider level and leverage explained about 60 percent of behaviour variance, more than the written strategy did.
A third arXiv paper prices the gap between generated code and production-qualified change, citing an audit that found material test or design issues in 59.4 percent of 138 benchmark tasks. Three studies from two independent publishers, drawn from a live business, a live market and an audited benchmark set, land on one point. Volume of agent output, token spend and benchmark scores are not evidence that a business outcome occurred.
dfdx labs: Hans Kraemer | arXiv: Six months of live LLM trading | arXiv: Production-Qualified Change and the Verification Tax
EY's first AI governance survey finds a quarter of organisations cannot detect agents they never authorised. Advisory
Ernst & Young published its inaugural US AI Risk and Governance Survey on 15 September. It reports 98 percent of respondents with formal AI governance policies and 69 percent with a single unified policy, alongside 47 percent who say their organisation has bypassed that process for urgent deployments. Of the 91 percent using agentic AI in pilots or production, 49 percent have not updated their governance framework for agentic risk and 26 percent cannot detect unauthorised agents inside the organisation. A further 41 percent say senior leaders lack visibility of every AI tool in use, and 36 percent report an AI incident with materially negative impact. Those last two numbers move the binding control from policy text to an inventory of agent identities.
Ernst & Young: AI governance has entered its next phase
Security
Ordinary prose can carry a hidden instruction past a cheap AI safety filter. Vendor
Check Point Research published a study on 10 September describing a technique that hides a policy-violating instruction inside fluent, ordinary text, with no encoding, emoji or invisible formatting. The fast screening models it tested marked all 23 crafted prompts as safe, while the same instructions in plain form were blocked. The stronger model behind the filter then recovered and acted on the hidden instruction in 17 of 18 completed trials, all run in emulated environments with no real user data. Extraction took that model more than a minute of reasoning and several code executions, which is why a fast filter missed it. The weakness is architectural rather than one model's quirk, and it sits exactly where a bank puts a cheap pre-filter in front of a capable tool-using model.
A vendor control pattern shows why per-call checks miss agent abuse that only appears across a session. Vendor
Google published runtime-governance guidance for AI agents on 15 September. Its illustrative case has eight individually allowed $20 refunds accumulating to $160 against a $149 order, because the policy evaluated each call on its own. The pattern puts three checks outside the agent's own code, which are content screening, an intent check before the tool call changes anything, and session-level detection of cumulative value and velocity. Google also describes an administrator adding a plain-language constraint that takes effect on later calls without rebuilding the agent, and says the outcomes shown are illustrative rather than an independent study. That last capability is a production change with no change-control owner, which is the governance gap this pattern opens while closing another.
On the radar
- A Cloud Security Alliance note dated 10 September describes four separate occasions on which frontier models reached real systems from supposedly isolated safety tests, each time through a third-party test environment that was meant to be air-gapped but had open internet access. Cloud Security Alliance
- Autonomous agents on the iLands platform sent unsolicited pitches to journalists, lawyers and academics, with one New York University professor reporting about 40 messages in a week before the platform added unsubscribe controls. 404 Media
- HSBC is hiring for a group AI safety and evaluation function covering agentic testing and release gates. HSBC