Pondero AI logo

Pondero AI

Archives
Log in
Subscribe
September 25, 2026

Pondero Brief: Cursor's bot now watches your deploy and drafts the revert

Pondero Brief - SEPTEMBER 25TH, 2026

Free Rollouts credits run out around Oct 3. Plus: OpenAI's agent hacked Medicare.
pondero. BRIEF · SEP 25

Cursor's bot now watches your deploy and drafts the revert

Also: OpenAI's agent wrote to a government database without authorization, and Gemini 4 is closer than expected.

Cursor shipped two bots on September 23 that work after the code is written. Rollouts watches the deploy and flags regressions per environment; Security Review scans every PR for exploitable bugs in 3.8 minutes. Enable Rollouts on Cursor Teams →

Also in today's brief

  • Cursor's bot now watches your deploy for regressions
  • OpenAI's agent broke into Medicare and went quiet for 30 days
  • Gemini 4 entered post-training ahead of schedule
  • Anthropic's new marketplace has 2,000 connectors and committed-spend
  • Amazon opened Seller Central to outside AI agents this week
 
Models & Releases

Gemini 4 is in post-training, and Gemini 3.5 Pro is not coming.

DeepMind chief Koray Kavukcuoglu said Google wants an early post-training version of Gemini 4 out "as soon as possible" and hopes for "much earlier" than the end of 2026, with Antigravity already running on it internally. The flagship teased over the summer is dead: Google "took a little bit of a step back" to focus on Flash, 9to5Google reports. Google's last flagship was Gemini 3 Pro in November 2025, so a full generation jump lands in a quarter where Anthropic and OpenAI have already refreshed. Signing a 12-month platform commitment in the next few weeks? Get a model-swap clause into it and hold two weeks of benchmarking time in reserve.

 
Tools & How-To

Claude Marketplace turns committed Anthropic spend into a purchasing budget.

Three doors opened September 23: connectors and plugins, more than 2,000 of them today; Claude-powered products from CrowdStrike, Cursor, Harvey, Legora, Lovable and Snowflake, bought against a portion of your committed Anthropic spend; and integrators from the Claude Partner Network. That middle door is the only part that changes a decision, because a tool you already wanted comes out of a budget line you have signed and security-reviewed, skipping the new vendor contract that usually gates it. Cursor is on the launch list, so the bots in today's top story are buyable with Claude credit.

 

Amazon let outside agents run a Seller Central account.

Amazon's selling partner plugin hands inventory, pricing, listings and analytics to Anthropic's Claude or Amazon's own Quick assistant, no Seller Central login needed. Connecting Claude takes about 60 seconds with no code, sellers scope the data and approve each action, and the beta is US-only, per GeekWire. Amazon also blocked Meta's Muse from shopping its store days earlier, so read this as a whitelist with two names on it. Sellers get real automation today; agent builders should plan for permission, not for an API.

 

LangChain's managed agents got user-scoped memory and a way out of Slack.

Managed Deep Agents 0.8 landed September 24 with user-owned credentials, memory scoped to the authenticated caller, HTTP webhook channels, Slack file transfer, and built-in Parallel web search, per LangChain. Scoped memory stops a shared agent dragging one person's context into a group thread; HTTP channels let it live in a customer portal instead of a Slack channel. Shelved a pilot over either? Reopen it.

 
Policy & Legal

An OpenAI agent broke into a Medicare portal, and OpenAI emailed a public tip-line 30 days later.

Researching public medical spending, an OpenAI agent hit authorization blocks on the Services Australia Medicare statistics portal and got through anyway on June 18, reaching aggregate statistics and internal file names, no patient records. It "found a way around those blocks, didn't accept 'no' for an answer," said Prime Minister Anthony Albanese. ABC's timeline is the damning part: OpenAI caught it on August 11 reviewing misaligned model activity in training, stayed quiet when Altman met Australia's defence minister on September 1, then on September 10 emailed the public inbox academics use to report bugs. Albanese says there will "obviously be legal consequences," per TechCrunch.

Copy the failure modes; both are in your stack. Detection came from OpenAI's own eval logs, not from anyone's alerting, because an agent routing around a 403 looks like ordinary traffic from the far side. Two questions for your agents this week: when one gets denied, does it retry down another path or stop and escalate? And who gets paged when the answer is the first one?

 
Quick Hits
• Alibaba put its agent stack on stage at Apsara 2026. AgentCore to build and run agents, an Agent Security Center beside it, and an Agent Context layer Alibaba says cuts token usage by up to 67 percent; Qwen 4 is in training. Worth a look if you serve Southeast Asia or need a cloud outside the US. Details →
• Google wants a coding agent that works with the internet switched off. The Gemma 4 Developer Agent competition asks entrants to post-train gemma-4-31b into an agent that fixes real GitHub issues offline on four L4 GPUs, closing December 2. Competition page →
 

How was today's brief?

★★★★★ Nailed it ★★★ Solid ★ Missed

Jonathan Hildebrandt Jonathan Hildebrandt
Co-founder and primary operator of Pondero. Writes the Pondero Brief.

Affiliate disclosure  ·  Unsubscribe  ·  Manage preferences

Pondero earns commissions on some links. This does not affect our editorial picks.

Don't miss what's next. Subscribe to Pondero AI:
← Newer Pondero Brief: Cursor's security bot hits 3.8-minute scans, and Copilot gets a meter Older → Pondero Brief: 950 agents, 21 hours, 210M tokens, one new enzyme
pondero.ai
Bluesky
LinkedIn
Twitter
LinkedIn
Powered by Buttondown, the easiest way to start and grow your newsletter.