Downstream News

Archives
Log in
Subscribe
August 19, 2026

Downstream — Wednesday, August 19, 2026

Downstream — Wednesday, August 19, 2026

27 stories, 54 corroborating sources. Deduplicated across vetted feeds and ranked for people building with agents.

1. Glean introduces Waldo agentic search model

Glean has developed Waldo, a specialized agentic search model that works in concert with frontier models to deliver end-to-end agentic outcomes, reducing latency by 50% and token usage by 25% while maintaining quality. Waldo is trained using a combination of direct preference optimization and reinforcement learning, and is designed to determine how much reasoning a task requires based on its own execution. The model will be deployed to customers soon.

Agents · Product · 4 sources

→ Latent Space — glean.com

Also covered by: AINews · AgentBrief — x.com · Simon Willison — softwaredoug.com

2. Glean Unveils Third-Generation AI Assistant

Glean introduces its third-generation AI Assistant with advanced personalization and agentic intelligence, and announces significant expansions to its Work AI platform, including new SDKs and MCP capabilities. The new Enterprise Graph underpins these innovations, enabling AI that truly understands the enterprise and its workflows. Glean Assistant now delivers complete outcomes tailored to each employee's way of working, without requiring advanced prompt engineering. The platform also features a new interactive workspace, Canvas, and allows employees to control their Assistant experience. Additionally, Glean is changing the way agents are built and deployed, making it easier for anyone to create and refine agents, and adding richer actions and MCP directory and host support.

Agents · Product · 4 sources

→ Latent Space — glean.com

Also covered by: AINews · AgentBrief — x.com

3. Agent Harnesses, Evals, and Production Feedback Loops

Miles v0.1, a new open-source RL framework, has been announced, and multiple developments in agent evaluation, search benchmarking, and harnesses have been reported, including LangSmith Tuned Evaluators and Managed Deep Agents, indicating a shift towards robust rollouts, CI, observability, and environment plumbing. Search benchmarking for agents is maturing, with Artificial Analysis launching its Search Index, and LangChain introducing LangSmith Tuned Evaluators, which claim better performance at lower cost.

Agents · Product · 2 sources

→ AINews — latent.space

Also covered by: AgentBrief

Also linked: @radixark announced Miles — x.com · Search Index — x.com · LangSmith Tuned Evaluators — x.com · +6 more

4. Frontier Model Cost and Open-Weights Popularity is Driving Demand for Model Routing

Glean's model routing helps control AI costs for organizations by selecting the most cost-effective model for each task, and its human feedback loop improves the routing system. The company has seen significant growth, reaching $300 million in annual recurring revenue, and is becoming a key player in the enterprise AI market. Glean's architecture includes a model called Waldo, which filters user queries and determines the best model to use, and the company is also seeing increased interest in open-weight models due to cost concerns.

Business · Product · 3 sources

→ Latent Space — latent.space

Also covered by: AgentBrief — x.com

5. Open Models: Qwen3.8-27B Momentum, GLM-5.3’s Post-Training Gains, and the Small-Model Debate

Qwen3.8-27B has become the top local model in Cline, with impressive benchmark results, while GLM-5.3 has been launched via API with significant gains in intelligence index and Elo rating, driven by stronger post-training techniques. The development of capable local models raises safety implications and highlights a shift towards useful, locally deployable models. GLM-5.3's gains suggest a shift in agentic capability scaling from parameter count to RL systems and environment quality.

Models · Product · 1 source

→ AINews — latent.space

Also linked: @kimmonismus calling it a “DeepSeek moment” — x.com · Alibaba Qwen celebrating it reaching #1 local model in Cline in four days — x.com · #7 on Artificial Analysis’ Agentic Index at 27B — x.com · +7 more

6. OpenAI’s Frontier RL Pause, Expanded Monitoring, and the Shift Toward “Pacing the Frontier”

OpenAI paused some frontier RL training for two weeks to strengthen monitoring, isolation, and red-teaming, and is holding its largest planned frontier RL run, with a focus on hardening security and alignment controls. The slowdown mainly affects farther-out releases, not models already near ship.

Safety · Product · 1 source

→ AINews — latent.space

Also linked: paused some frontier RL training for two weeks — x.com · capabilities were outpacing safety/alignment readiness — x.com · confidence in safety will increasingly set the pace of frontier scaling — x.com · +3 more

7. Research Notes: Multi-Agent Coordination, Training Variance, and Public AI Usage Measurement

The Public AI Observatory is a new effort to measure real AI assistant usage, with 24,521 consented conversations and 52 models analyzed, while separate research highlights key findings on multi-agent teams and training variance, including the impact of task structure on communication topology and the emergence of specification gaming in agent collectives. The observatory aims to provide public-interest observability for AI usage patterns, independent of vendor reporting.

Agents · Product · 1 source

→ AINews — latent.space

Also linked: @omarsar0 — x.com · @sfrei_ — x.com · Public AI Observatory — x.com

8. Inference and Systems Infra: Mojo Open Source, TensorRT Connect, Cursor’s Git Storage, and Faster Decoding

Modular has open-sourced Mojo, positioning it as a portability layer across accelerators, while NVIDIA launched TensorRT Model Connect for direct model conversion and deployment, and Cursor published a retrospective on Git hosting at scale. Additionally, several companies announced advancements in on-device inference and datacenter accelerators, including DFlash 2 and Cerebras CS-4, highlighting the increasing importance of inference speed.

On-device · Product · 2 sources

→ AINews — latent.space

Also linked: the company formally open-sourcing Mojo — x.com · Qualcomm datacenter AI accelerators — x.com · TensorRT Model Connect in public preview — x.com · +5 more

9. DeepSeek Harness breaks record with 20k stars in one hour

The DeepSeek Harness repository has achieved a record-breaking 20k stars in approximately one hour, surpassing previous records set by AutoGPT and Grok-1

Agents · Product · 3 sources

→ Sarah Guo — x.com

Also covered by: AINews — x.com

10. Glean introduces model choice for Assistant conversations

Glean now allows users to select a specific AI model for each Assistant conversation, with options including GPT, Claude, and Gemini models, and also offers an Auto option that automatically selects the best model based on internal evaluations and live usage data. Admins can control which models are available for their users and set default models for their organization.

Coding · Product · 1 source

→ Latent Space — docs.glean.com

11. Top tweets (by engagement)

Anthropic's Claude autonomously designed protein binders for 14 out of 15 targets, and Claude gained Gmail and Google Drive actions, while other AI developments and updates were also announced

Agents · Product · 1 source

→ AINews — latent.space

Also linked: @sama on pausing frontier RL training pending stronger safety/alignment standards — x.com · @AnthropicAI on Claude autonomously designing protein binders for 14/15 targets — x.com · @OpenAI detailing the two-week pause and new security/monitoring controls — x.com · +3 more

12. Perplexity Computer adds DeepSeek V4 Pro

DeepSeek V4 Pro is now available in Perplexity Computer, offering a cost-effective solution with a 62% cheaper cost-performance ratio compared to the next model, and it scored 0.359 at $0.75 per task on WANDR

Models · Internals · 4 sources

→ Perplexity — x.com

Also covered by: AINews · AgentBrief — x.com · @InsiderPhD

13. AI struggles with instruction following

AI models are experiencing issues with following instructions, a problem that may not be rapidly solved, similar to the challenge of hallucinations

Safety · Product · 5 sources

→ Gary Marcus — x.com

Also covered by: AINews · AgentBrief — x.com · Robert Scoble · Bojan Tunguz

14. Anthropic Claude Code usage grows to 4% of GitHub

Anthropic's Claude Code has seen significant growth, now accounting for around 4% of GitHub code, and its usage is discussed in the context of AI engineering and semi-analysis work. The episode also touches on Memory Mania and its potential impact on users.

Coding · Product · 2 sources

→ AINews — latent.space

15. Coding agents streamline app development

Coding agents have improved the development process by reducing barriers between different teams, allowing for faster creation of working apps

Coding · Product · 2 sources

→ Robert Scoble — x.com

Also covered by: AgentBrief

16. Glean outperforms Claude Cowork in cost benchmark

Glean's harness and routing capabilities were benchmarked against Claude Cowork, showing a 4x cost advantage due to lower token volume and cheaper rates. Glean achieved this through model family routing, model tier routing, and better context handling, consuming fewer tokens per query. More results will be presented at Glean:GO!

Agents · Product · 2 sources

→ Latent Space — x.com

Also covered by: AINews

17. OJO simplifies tool creation process

OJO streamlines the process of turning ideas into browser-accessible tools by handling the layer above coding agents, reducing the need for multiple handoffs and extensive development work

Coding · Product · 2 sources

→ Robert Scoble — x.com

Also covered by: AgentBrief

18. Zai releases GLM 5.3 with new vulnerability detection

Zai released GLM 5.3, which was evaluated by Semgrep for its vulnerability detection capabilities, and the results show potential for open weight models in cyber security tasks

AI security · Product · 2 sources

→ @InsiderPhD — x.com

19. Zillow adopts Glean for AI-forward culture

Zillow implemented Glean to unify its fragmented data landscape, enabling employees to quickly find information and deploy specialized agents across critical workflows, resulting in improved onboarding, customer focus, and engineering acceleration. Glean's integration with MCP servers and AI coding assistants like Claude Code also accelerated project initiation and raised coding standards.

Agents · Product · 2 sources

→ Latent Space — glean.com

Also covered by: AINews

20. Booking.com adopts Glean AI platform

Booking.com implemented Glean's AI and search platform to improve information access and productivity, resulting in reduced video script creation time and faster IT ticket resolution. The company also integrated AI into its strategy and workflows, adopting Glean as its first company-wide AI platform

Business · Big picture · 1 source

→ Latent Space — glean.com

21. Glean updates Model Hub configuration

Glean has updated its Model Hub to allow configuration of available models and selection of models for workflows, with options for Glean Universal Model Key and Customer Key deployments. The update includes best practices for model selection and management.

Coding · Product · 1 source

→ Latent Space — docs.glean.com

22. Researchers unify 3DGS with Gaussian Splats

A new paper presents a framework that unifies 3DGS captured assets using Gaussian Splats, potentially improving physics understanding in robots. This could lead to advancements in areas like table tennis-playing robots.

Robotics · Product · 1 source

→ Robert Scoble — x.com

23. Rox builds sales agents on top of CRM backend

Rox developed agents that automate research, outreach, deals, and renewals on top of a CRM backend, with successful implementations on MongoDB and Together

Agents · Product · 1 source

→ Robert Scoble — x.com

24. Semgrep releases GLM 53

Semgrep's GLM 53 delivers Opus 48 level cybersecurity results at a lower cost, as analyzed by the company's colleagues

AI security · Product · 3 sources

→ @InsiderPhD — x.com

Also covered by: AINews · AgentBrief — x.com

25. CC AI agent opens waitlist in Australia

CC AI productivity agent in Gmail opens waitlist in Australia and New Zealand, and expands availability in the US and Canada, rolling out invitations to waitlisted users

Coding · Product · 1 source

→ Google Labs — x.com

26. Memory prices surge 500% in 12 months

Memory prices have increased by 500% in the past 12 months, with 128GB DDR5 kits now costing $3,399, due to high demand from AI datacenter buildouts and limited supply. This price surge is affecting not only RAM but also other components like hard drives and SSDs.

Business · Big picture · 1 source

→ AINews — tomshardware.com

27. Perplexity Computer integrates email functionality

Perplexity's Computer now allows users to interact with it via email, enabling send, forward, or cc actions on any thread, with tasks running as normal sessions and maintaining an audit trail

Agents · Product · 1 source

→ Perplexity — x.com

From Around the Web

1. Cerebras launches CS-4, 30x faster inference than GPUs

Cerebras introduced the CS-4, a rack-scale AI solution that delivers up to 30x faster inference compared to GPUs, with enhanced economics and simplified deployment, featuring three WSE-3 Turbo per system and a modular Nexus Rack-Scale Platform. The CS-4 solution shifts the inference Pareto frontier, delivering up to 10x more throughput per watt than CS-3 while generating tokens up to 30x faster than production GPU systems.

On-device · Product · 3 sources

→ hackernews — cerebras.ai

Also covered by: AINews · AgentBrief — x.com

2. The Mojo language (by Modular, now Qualcomm) is now open-source

Modular has made its Mojo language and Modular Cloud platform open source and publicly available, supporting various hardware accelerators including AWS Trainium, Google TPUs, and Qualcomm Cloud AI 100, with native Windows support coming soon. The Modular Platform is now production-ready, serving billions of tokens per minute and powering real enterprise deployments, with flagship customers like MiniMax. Modular is also opening up its MAX licensing model and expanding source access to build an industry alliance program.

Coding · Product · 2 sources

→ hackernews — modular.com

Also covered by: AgentBrief

3. AI usage patterns in software teams

Linear's report shows a significant increase in AI adoption and pull requests among its users, with coding agents driving most of the acceleration. The report also highlights the blurring of roles, with senior leaders and non-engineers committing code. However, the increased output has not led to time savings, with teams working more, not less.

Coding · Product · 2 sources

→ hackernews — linear.app

Also covered by: AgentBrief

4. fx releases v0.0.3 coding agent harness

fx is a minimalistic, open-source coding agent harness and CLI written in Zig, optimized for research and embeddability, with a focus on performance and minimalism. It features a small binary size, instant installation, and embedding capabilities, making it suitable for resource-constrained environments and agent sandboxes.

Coding · Product · 2 sources

→ hackernews — fx.sh

Also covered by: AINews

5. Palomar: A registry of Lean verified mathematics

The Palomar registry, an initiative for verified mathematics, is now open for submissions of Lean proofs, providing a platform for formalizing and verifying mathematical results. The registry checks submissions for correctness and adherence to best practices, and welcomes both human-generated and AI-generated proofs.

Research · Internals · 1 source

→ hackernews — terrytao.wordpress.com

6. GLM-5.3 Artificial Analysis Benchmarks

GLM-5.3 (max) is a reasoning model that scores 60 on the Artificial Analysis Intelligence Index, outperforming the median score of 35, and is priced at $1.40 per 1M input tokens and $4.40 per 1M output tokens. The model has a context window of 1M tokens and generates output at 74 tokens per second.

Models · Product · 2 sources

→ hackernews — artificialanalysis.ai

Also covered by: AINews

7. Bun 1.4 Rust rewrite is not looking good

The Bun 1.4 rewrite, which heavily utilizes AI-powered tools like Anthropic's Claude, has been plagued by delays and criticism from the community, with over 5,000 open pull requests and concerns about code quality and safety. The project's creator, Jarred, has faced backlash for making repeated promises of imminent release, only to delay again. The rewrite has also raised questions about the effectiveness of AI-assisted coding and the decision to switch from Zig to Rust.

Coding · Product · 2 sources

→ hackernews — tipiirai.com

Also covered by: Simon Willison


Read this digest on the web · Archive · RSS

Downstream points you at the primary source; it does not replace it.

Don't miss what's next. Subscribe to Downstream News:
← Newer Downstream — Friday, August 21, 2026 Older → Downstream — Tuesday, August 18, 2026
pablooliva.de
Powered by Buttondown, the easiest way to start and grow your newsletter.