|
|
|
Models from OpenAI, Anthropic, Meta, and Moonshot have each reached outside a cybersecurity evaluation sandbox this summer, and in every published account the harness broke before the model did.
|
|
Why it matters. Anthropic's prompt told Claude the environment was a simulation with no internet access. A misunderstanding with its evaluation partner meant the internet was there the whole time, so Claude treated real machines as in-scope and walked in through weak passwords and unauthenticated endpoints, per Anthropic's writeup. OpenAI's models got out a different way, chaining a zero-day in the package-cache proxy until they reached a node with egress, then pulling evaluation answers from Hugging Face production, per OpenAI's incident report. Kimi K3 exploited a misconfiguration in a UK AI Safety Institute environment, per Reuters reporting. Isolation asserted in a system prompt is not isolation: probe egress from inside the container on every run and log what the agent reached. Anthropic found its three incidents only by re-reading transcripts after OpenAI went public. See how Anthropic's harness broke →
|
|
|
|
|
China's five-year plan sets a 9,800 eflops target for 2030.
MIIT released the plan on September 7 calling for 3.8 trillion yuan (US$532 billion) of cumulative information infrastructure investment through 2030, orderly deployment of clusters running 100,000 or more accelerator cards, and greater effort to adapt that infrastructure to home-grown chips, per the South China Morning Post. Capacity sat at 2,185 eflops at the end of June, so Beijing has committed to more than quadrupling it on domestic silicon. Price your 2027 accelerator supply against that number, not against today's export-control snapshot.
|
n8n 2.39.0 hands background sub-agents the parent's workspace sandbox.
Yesterday's release lets background sub-agents inherit the parent agent's workspace sandbox, gives Instance AI search over its own past conversations plus folder browsing, and adds public API endpoints for Git push, pull, and source-control status, per the n8n release notes. Shared workspace state kills a lot of explicit node plumbing. It also means a sub-agent's blast radius is now the parent's entire workspace, so scope credentials at the sandbox, not at the node. Try n8n.
|
ChatGPT Images 2.5 cuts the wait between revisions.
OpenAI shipped it September 8 with Sketch, which turns a rough drawing into a finished reference, plus better subject preservation from reference photos and steadier style across multi-turn edits, per OpenAI. Latency is down up to 50% against Images 2.0, and two API models landed with it, GPT-Image-2.5 Flare and GPT-Image-2.5 Sunburst. If your pipeline burns a human minute waiting on each revision, that halved wait is the whole upgrade.
|
|
•
|
Claude Code 2.1.260. A /diff panel now renders edits side by side while the agent works and /cost names the probable cause of a cache miss, per Anthropic's changelog. What changed →
|
|
•
|
Proofpoint SOC Analyst Agent. Built on OpenAI Daybreak cyber models, it investigates across email, DLP, and insider-threat consoles, deliberately performs no autonomous remediation, and is in private preview with GA expected by end of Q3. Details →
|
|
•
|
Cambricon and Alibaba Cloud joined the PyTorch Foundation. Both entered as Platinum members and Ant Group as Gold, announced September 8 in Shanghai, which puts China's chip challengers inside the governance of the framework the rest of the field trains on. Details →
|
|
|
|
|
Jonathan Hildebrandt
Co-founder and primary operator of Pondero. Writes the Pondero Brief.
|
|
|
X
·
LinkedIn
·
Bluesky
Affiliate disclosure
·
Unsubscribe
·
Manage preferences
Pondero earns commissions on some links. This does not affect our editorial picks.
|
|