Weekly dispatch

The Newsletter

One short email a week — the AI ideas, tools, and questions worth your attention.

Want a curated newsletter?

Tell me the topics, teams, or questions you'd like tracked each week and I'll put together a dispatch shaped around them.

Zero spam, zero funny business — your email is only used to reply about your request. Bail out any time; we won't take it personally (much).

Sep 14, 2026 – Sep 20, 2026

Weekly AI News Digest

Coverage window: Monday, September 14, 2026 – Sunday, September 20, 2026
Publish date: Monday, September 21, 2026


1. Open weights win the token race; Anthropic still owns the bill

Category: AI models / open source / industry usage
Source: The New Stack
Date: September 18, 2026

Vercel’s September AI Gateway report (covering August traffic) puts a hard number on the open-weight surge: open-weight models handled 56% of all tokens routed through the gateway — the first month they crossed a majority — up from 7% in December 2025, 13% by April, and 36% in July. CEO Guillermo Rauch flagged August 22 as a record day at 62% open-weight share, arguing enterprise adoption is still early and that harnesses, CLIs, IDEs, and SDKs still need to become model-agnostic.

Token majority is not dollar majority. Open-weight traffic (DeepSeek, Moonshot, Z.ai, and peers) was only about 14¢ of every estimated dollar through the gateway, while Anthropic took 64¢ of every dollar — and has not fallen below 61% of spend in any month since December 2025. Inside that Anthropic share, customers rotated: Fable 5’s spend share dropped as cheaper Opus 5 climbed, evidence that “lab loyalty” follows model profile more than brand. Average price per token on the gateway fell 23.2% in August (third straight monthly decline). The New Stack also notes the parallel OpenRouter signal: open-weight models were ~60% of US-originating token consumption in August, largely Chinese-developed models served from US providers.

Why it matters: This is the clearest production usage chart of the week. Open weights are winning volume; frontier labs (especially Anthropic) are still winning the wallet — and the gap is exactly where gateway, router, and harness vendors are pitching.

Source: The New Stack — Open-weight tokens vs Anthropic spend


2. Better harnesses beat better models: Zed, Anthropic, and OpenRouter’s week

Category: AI engineering / developer tools / agentic platforms
Source: The New Stack
Date: September 19, 2026

The New Stack’s weekly read framed five hot stories as one thesis: the product is the harness — context, tools, routing, and verification around the model — not the next checkpoint weights. Inference keeps getting cheaper (see Vercel’s 23.2% token-price drop); capital is moving to the software that turns a model into something teams can ship.

Three product moves carried the argument. Zed opened Delta in public beta: shared threads instead of pull requests, with DeltaDB recording edit-level changes so the agent conversation stays attached to the code. CEO Nathan Sobo’s line — “everyone is in a race to replace GitHub right now” — matched the load story (GitHub at 2.9B monthly commits in August after 1.4B in April). Zed says 33 of its own teammates landed 570 changes on Delta’s main branch without a single PR. Anthropic began folding Claude Chat and Cowork into one interface so users stop deciding which mode a task belongs in. OpenRouter made US in-region routing generally available for business/enterprise — decrypt, process, and serve inside the US, or reject the request — so buyers can separate model origin from where data is processed as open-weight Chinese models dominate US token volume.

Why it matters: After last week’s capital (Mistral) and enterprise control-plane (Salesforce) stories, this week’s AI-engineering headline is collaboration and routing UX: threads over PRs, one chat surface over mode picking, and sovereignty knobs on the router.

Source: The New Stack — Harness economics (Zed / Anthropic / OpenRouter)


3. Video roundup: Enterprise AI infra reality check + agent memory / inference engineering

Category: Platform engineering / AI engineering / cloud-native AI (video)
Sources: YouTube — @PlatformEngineering, @aiDotEngineer, @cncf
Dates: September 16–19, 2026

Primary pick — @PlatformEngineering (streamed September 17; uploaded September 18): Enterprise AI infrastructure in 2026: A reality check for platform engineers — Dan Ciruli, Kelsey Hightower, and Luca Galante on folding inference, fine-tuning, and agents into existing heterogeneous platforms (VMs + containers + legacy data) rather than rebuilding for greenfield training. The session’s ADP (Agentic Development Platform) framing is the PE counterpart to this week’s written harness thesis.

Also worth watching this week:

Why it matters: Written coverage is tokens-vs-dollars and harness productization; the video circuit is platform teams absorbing inference/agents into existing infra, plus a dense AI Engineer inference + memory week.


4. The Human Guide to AI — still waiting on the next chapter

Category: AI literacy / human stories of AI
Source: Medium — The Human Guide to AI
Date: Latest publication post remains July 17, 2026 (no new article this coverage window)

The publication has not shipped a new piece since The Genesis of Mind (July 17). That article — and the series opener Who is AI? — were already shortlisted in prior digests, so this edition does not rehash either narrative.

This week’s Human Guide slot is a publication watch: the latest available article remains Genesis of Mind. Readers new to the series should start with the publication homepage (Who is AI? → Genesis of Mind). We will feature a full write-up as soon as a newer post lands.

Why it matters: The required Human Guide source is checked without recycling prior shortlists — while open-weight usage charts and harness product news dominate the firehose.

Source: Medium — The Human Guide to AI · latest post: The Genesis of Mind


5. When the agent fails, debug the runtime — Nvidia’s SAFE exchange + OpenShell

Category: Cloud-native agentic infra / AI engineering / observability
Source: The New Stack
Date: September 20, 2026

Nvidia VP of Product Adel el Hallak told The New Stack that production agents fail without looking like conventional software failures — they keep running while “getting creative,” and top coding agents still miss on >60% of real-codebase tasks. The fix is not “more logs of inputs and outputs”; it is reasoning traces, tool choices, stuck points, and approach changes — often by replaying execution. What looks like a model bug may be harness or runtime.

Nvidia’s stack framing: model (intelligence) / harness (orchestration) / runtime (governance). The non-negotiable piece in its reference architectures is OpenShell (under NemoClaw) for sandboxing, policy, and visibility — “change the harness or the model; keep the secure open runtime.” Industry sharing layer: the Secure Agent Findings Exchange (SAFE), backed by ~140 companies, aims to circulate agent failure findings the way vulnerability disclosure works for traditional software. Nvidia’s NOAH research underscored the harness lever: same model, different harness, different outcomes.

Why it matters: Complements this week’s “buy the harness” product news with the operations question: when agents ship, who owns the failure taxonomy — and can the industry share it?

Source: The New Stack — Nvidia agent debugging / SAFE


Sep 7, 2026 – Sep 13, 2026

Weekly AI News Digest

Coverage window: Monday, September 7, 2026 – Sunday, September 13, 2026
Publish date: Monday, September 14, 2026


1. Mistral’s $3.5B bet: open weights are not enough without the stack

Category: AI models / open source / industry infrastructure
Source: The New Stack
Date: September 10, 2026

Mistral closed a €3 billion Series D (~$3.5B), lifting post-money valuation past €21 billion, and framed the raise as fuel for frontier research, training compute, and infrastructure — not just another open-weight model drop. The New Stack’s read is structural: downloadable weights alone do not break concentration when chips, training runs, and high-volume serving still sit with a short list of labs and cloud providers.

That tension is explicit in the industry argument the piece tracks. Anthropic CEO Dario Amodei’s recent line — open weights “are nowhere near a sufficient solution because they simply shift the concentration somewhat to those with the most compute and chips” — is the foil. Mistral’s answer is to build more of the stack: open-weight models plus the capacity to train and serve them at frontier scale. The company’s own positioning in the funding news leans on being an AI firm that ships across model and infrastructure layers rather than treating openness as a weights-only virtue signal.

Why it matters: After a summer of harness FinOps, router wars, and open-weight releases, this week’s model story is capital and compute. Open-source credibility increasingly means owning enough of the training/serving plane that “open” is an operational option, not just a license checkbox.

Source: The New Stack — Mistral $3.5B / open infrastructure


2. Salesforce’s Enterprise AI Harness: six tools, one control plane

Category: Platform engineering / enterprise agentic / AI engineering
Source: The New Stack
Date: September 10, 2026

Salesforce launched its Enterprise AI Harness as a formal packaging of what it calls six trusted platform capabilities under a shared AI control plane: Data 360, Informatica, MuleSoft + Agent Fabric, Tableau, Agentforce, and Salesforce Guardian, on top of the Salesforce platform itself. The pitch is blunt about “AI leakage”: CRM, ERP, field service, and support each know a slice of a customer order, so siloed agents keep automating incomplete context.

Rohan Kumar (president & chief platform and engineering officer) frames differentiation away from model choice: “The Agentic Enterprise won’t be defined by which model a company chooses… what will differentiate an enterprise is the trusted, proprietary context it brings to that intelligence.” Teams can take the six together or à la carte, and mix Salesforce with third-party pieces. Subsystems such as Agentforce Vibes already lean on specialized execution harnesses (Mastra, Claude Agent SDK); this release is the company-wide control-plane story — context, agency, action, governance, security, and models aligned so policies and workflows are reusable across agents.

Why it matters: This is the enterprise PE / IDP analogue of last month’s harness discourse. The durable product is not “pick Claude or Astra,” it is a composable control plane that keeps proprietary customer context and governance attached when agents cross system boundaries.

Source: The New Stack — Salesforce Enterprise AI Harness


3. Video roundup: Booking.com’s GenAI platform + Block ACP + Kubeflow agent SQL

Category: AI engineering / platform engineering / cloud-native AI (video)
Sources: YouTube — @PlatformEngineering, @aiDotEngineer, @cncf
Dates: September 9–10, 2026

Primary pick — @PlatformEngineering (September 10): Beyond access: Building the GenAI platform behind 3000+ developers at Booking.com. A concrete PlatformCon London talk on moving past “give everyone a chatbot” to a GenAI platform that can serve thousands of developers — the PE scale story of the week.

Also worth watching this week:

Why it matters: Written coverage this week is capital + enterprise harness + long-running agent APIs; the video circuit is platform scale (Booking.com), agent client protocols and skills catalogs, and Kubernetes-native constraints for agent tool use.


4. The Human Guide to AI — still waiting on the next chapter

Category: AI literacy / human stories of AI
Source: Medium — The Human Guide to AI
Date: Latest publication post remains July 17, 2026 (no new article this coverage window)

The publication has not shipped a new piece since The Genesis of Mind (July 17). That article — and the series opener Who is AI? — were already shortlisted in prior digests, so this edition does not rehash either narrative.

This week’s Human Guide slot is a publication watch: the latest available article remains Genesis of Mind. Readers new to the series should start with the publication homepage (Who is AI? → Genesis of Mind). We will feature a full write-up as soon as a newer post lands.

Why it matters: The required Human Guide source is checked without recycling prior shortlists — while Mistral’s stack thesis and Salesforce’s control-plane packaging dominate the enterprise firehose.

Source: Medium — The Human Guide to AI · latest post: The Genesis of Mind


5. OpenAI Agents API: Codex’s long-running harness goes public beta

Category: Cloud-native agentic infra / AI engineering / FinOps
Source: The New Stack
Date: September 11, 2026

OpenAI opened the Agents API in public beta — the backend behind Codex — so developers can run unattended agents for hours or days without rebuilding job tracking, execution surfaces, context compaction, or on-demand tool/subagent fan-out. Work can land in OpenAI’s sandbox or on customer-controlled infrastructure; developers pay for models, tools, and hosted compute, while orchestration sits outside that bill.

The usage math is the headline. An OpenAI research report (September 6) said the research org was logging 3.1 agent-workdays per human workday by mid-August; the median researcher (by agent usage) spent >$600/day on inference at API prices, and the 90th percentile exceeded $7,000/day. Same-day irony: OpenAI paused new $200 Pro sign-ups after GPT-6 Astra demand strained capacity. Agents API has separate rate limits, but the capacity story and the “make long-running agents easier” story arrived together.

Why it matters: Last week’s Astra launch/harness A/B was about frontier model packaging; this week’s OpenAI news is the productized agent runtime — and the inference bill that appears when agents outwork their humans.

Source: The New Stack — OpenAI Agents API / compute


Aug 31, 2026 – Sep 6, 2026

Weekly AI News Digest

Coverage window: Monday, August 31, 2026 – Sunday, September 6, 2026
Publish date: Monday, September 7, 2026


1. GPT-6 Astra’s “AGI era” — and the harness that scored 98.6%

Category: AI models / AI engineering / agent harnesses
Source: The New Stack
Date: September 3–4, 2026

OpenAI launched GPT-6 Astra on Thursday — its largest training run yet (100,000+ GPUs at Stargate Texas), priced at $10 / $50 per million input/output tokens, with rollout starting via Daybreak enterprise customers before Plus/Pro/Business/API. President Greg Brockman framed the moment as the start of an “AGI era,” while The New Stack’s launch write-up notes Astra does not clearly lead the coding pack: DeepSWE ~74.1% sits inside a tight cluster with Muse Spark 1.3, Gemini 3.8 Flash, and Claude Opus 5.

The sharper story landed a day later. ARC Prize ran the same model through its standard harness and got 62.7% on ARC-AGI-3; OpenAI’s Provider Adapter (opaque reasoning-state continuity + compaction) scored 98.6% — a 36-point swing, with the adapter run also cheaper ($17.3k vs $26.1k at max effort). Across shared solved tasks, the adapter used ~49% fewer tokens and ran ~3.66× faster. OpenAI documents the Responses API primitives; what you still can’t buy is the assembled system that posted the headline number.

Why it matters: Frontier coding scores are converging; the durable wedge is harness engineering — memory, resume, compaction, and the software around the model. Astra is the model of the week; the ARC Prize A/B is the platform lesson.

Sources: The New Stack — GPT-6 Astra launch · The New Stack — ARC-AGI asterisk · The New Stack — Astra harness vs ARC Prize


2. Coder Agent Relay: Cursor cloud agents on regulated infra

Category: Platform engineering / enterprise agentic coding
Source: The New Stack
Date: September 4, 2026

Coder launched Agent Relay with SpaceXAI (Cursor) as launch partner: coding agents keep Cursor’s cloud inference/planning loop, while tool execution and repo checkout stay in self-hosted Coder workspaces on the customer network. The pitch targets banking, life sciences, defense, aerospace, and government — places where “block Cursor outright” has been a common security outcome, not because the product is broken, but because the deployment model was wrong.

Coder CEO Rob Whiteley frames the FinOps red herring bluntly: “1% of my engineers are responsible for 40% of my token spend” — the real gap is uneven agent adoption, not tokenmaxxing. Agent environments are sandboxed, ephemeral, and scoped to one task; prompt-injection paths that would reach unauthorized resources are blocked at the environment layer; every run logs what was accessed, executed, changed, or denied. Private preview with design partners.

Why it matters: This is the PE / IDP story of the week — golden paths so regulated developers can use the agent UX they want without becoming the systems integrator, and without shipping source/secrets into a vendor cloud.

Source: The New Stack — Coder Agent Relay / SpaceXAI


3. Video roundup: incident response + million-token agents + OTel for AI

Category: AI engineering / platform engineering / cloud-native AI (video)
Sources: YouTube — @PlatformEngineering, @aiDotEngineer, @cncf
Dates: August 31 – September 4, 2026

Primary pick — @PlatformEngineering (September 4): What does good incident response look like end-to-end? Platform, AI & more. Fresh PE-channel framing for platform + AI ops while agents raise the stakes of detection, response, and control.

Also worth watching this week:

Why it matters: Written coverage this week is Astra + enterprise agent deployment; the video circuit is PE incident craft, long-context / identity-tethered agents, and cloud-native observability for agentic systems.


4. The Human Guide to AI — still waiting on the next chapter

Category: AI literacy / human stories of AI
Source: Medium — The Human Guide to AI
Date: Latest publication post remains July 17, 2026 (no new article this coverage window)

The publication has not shipped a new piece since The Genesis of Mind (July 17). That article — and the series opener Who is AI? — were already shortlisted in prior digests, so this edition does not rehash either narrative.

This week’s Human Guide slot is a publication watch: the latest available article remains Genesis of Mind. Readers new to the series should start with the publication homepage (Who is AI? → Genesis of Mind). We will feature a full write-up as soon as a newer post lands.

Why it matters: The required Human Guide source is checked without recycling prior shortlists — while Astra’s AGI rhetoric and regulated agent deployment dominate the engineering firehose.

Source: Medium — The Human Guide to AI · latest post: The Genesis of Mind


5. Nvidia PAIR: idle Macs and PCs as a home agent inference mesh

Category: Open-source / local agentic infrastructure / AI models
Source: The New Stack
Date: September 3, 2026

Nvidia shipped Personal AI Router (PAIR) — open-source software that discovers idle Macs/PCs on a LAN (mDNS), routes agent model requests through existing Ollama or LM Studio installs, and returns results through one familiar local endpoint. It is explicitly not a new inference engine and does not pool VRAM or shard a single request across machines: each request runs start-to-finish on one eligible node. The design target is speeding agents (NemoClaw, OpenClaw, Hermes, and peers) by fanning work across subagents in parallel on spare hardware.

Why it matters: While hyperscalers sell hosted harnesses, PAIR is a consumer/edge counter-move — local, open, agent-shaped routing that turns household silicon into a personal inference mesh without replaying last week’s Hugging Face M&A story as the lead.

Source: The New Stack — Nvidia PAIR local inference


Aug 24, 2026 – Aug 30, 2026

Weekly AI News Digest

Coverage window: Monday, August 24, 2026 – Sunday, August 30, 2026
Publish date: Monday, August 31, 2026


1. Nvidia’s $12.9B Hugging Face bet — owning where open models live

Category: AI models / open source / industry trends
Source: The New Stack
Date: August 27–28, 2026

Nvidia reportedly agreed to buy Hugging Face for $12.9 billion — about twelve days of Nvidia’s latest quarterly revenue — in a deal The Information first flagged Wednesday. The strategic rhyme is Microsoft/GitHub: follow the developers. Hugging Face is where open weights land, where fine-tunes ship, and where teams figure out how to run a model. Nvidia already publishes Nemotron open models; this purchase is about the distribution surface, not a shortage of weights.

The New Stack’s follow-up frames the open-source problem: Hugging Face’s value rides on hardware neutrality (Optimum for Nvidia and AMD/Intel/AWS). Owning the Hub gives Nvidia more room to pull NIM and CUDA-optimized paths into the default deploy experience. Competing silicon does not have to vanish to become less appealing if the Nvidia path takes fewer steps. Meanwhile the open-model lane keeps widening — Ollama’s Claude Desktop bridge for Qwen/DeepSeek/Kimi, laptop-sized open coding models, and CFOs noticing frontier API bills.

Why it matters: Last week’s opening shortlist was routers and software factories. This week’s industry lead is who owns the open-model Hub — and whether that Hub stays a neutral runway or becomes a chipmaker’s on-ramp.

Sources: The New Stack — Nvidia open models on its chips · The New Stack — Hugging Face deal’s open-source problem


2. Microsoft Agent Lightning v1.0: train through the production harness

Category: Platform engineering / AI engineering / agentic RL
Source: The New Stack
Date: August 26, 2026

Microsoft Research’s Agent Lightning v1.0 (GitHub tag August 16; covered for platform engineers this week) attacks a structural mismatch: traditional agentic RL lets the training engine own the interaction loop, while production agents live inside a different harness. Lightning flips it — the harness owns context construction, tool execution, and the agent–environment loop; the trainer only observes LLM request/response pairs across a service boundary.

On “modest compute” (6K training examples), RL via Lightning lifts Qwen3.5-9B on SWE-bench Verified from 41.8% → 56.4% (+14.6 points). Microsoft also ships a data-cleaning pipeline and reproducible scripts on open datasets/models so coding-agent RL is not a research-only artifact. External commentary to The New Stack stresses the real win: less train–serve mismatch when tool protocols, context policy, and recovery behavior stay identical from training to production.

Why it matters: This is the platform-engineering story of the week closing — IDP/harness teams can keep the production agent architecture as an asset instead of a training-time liability.

Source: The New Stack — Microsoft Agent Lightning harness


3. Video roundup: Backstage IDPs + agents as distributed systems + KubeCon Japan

Category: AI engineering / platform engineering / cloud-native AI (video)
Sources: YouTube — @PlatformEngineering, @aiDotEngineer, @cncf
Dates: August 26 – August 30, 2026

Primary pick — @PlatformEngineering (August 28): What it really takes to build an Internal Developer Platform with Backstage. Fresh PE-channel grounding for IDP builders while agents become first-class platform consumers.

Also worth watching this week:

Why it matters: Written coverage this week is Hub ownership + harnessed RL; the video circuit is IDP craft, SRE lessons for agents, and the cloud-native community calendar.


4. The Human Guide to AI — still waiting on the next chapter

Category: AI literacy / human stories of AI
Source: Medium — The Human Guide to AI
Date: Latest publication post remains July 17, 2026 (no new article this coverage window)

The publication has not shipped a new piece since The Genesis of Mind (July 17). That article — and the series opener Who is AI? — were already shortlisted in prior digests, so this edition does not rehash either narrative.

This week’s Human Guide slot is a publication watch: the latest available article remains Genesis of Mind. Readers new to the series should start with the publication homepage (Who is AI? → Genesis of Mind). We will feature a full write-up as soon as a newer post lands.

Why it matters: The required Human Guide source is checked without recycling prior shortlists — and the literacy gap remains while Hub M&A and harness RL dominate the engineering firehose.

Source: Medium — The Human Guide to AI · latest post: The Genesis of Mind


5. Same model, different harness: token use varied 70×

Category: AI engineering / FinOps / agent harness research
Source: The New Stack
Date: August 27, 2026

Three recent benchmarking efforts argue that when you price a coding agent, the harness may matter as much as the model. A June independent suite ran Aider, Claude Code, OpenClaw, and peers through OpenRouter on identical tasks: tokens per solved task ranged from roughly 3,500 (Aider architect mode) to 292,000 (OpenClaw) — about 70× — and the ordering barely moved across two unrelated models. Composio’s August harness bake-off on DeepSeek V4 Flash put successful-task cost from $0.028 (Pi Agent) to $0.195 (Claude Code); DeepAgents matched Claude Code’s pass rate at about a quarter of the cost. Artificial Analysis adds statistical weight with a multi-benchmark coding-agent index holding the model fixed while swapping Claude Code / Cursor CLI / Opencode.

The mechanism is the startup tax: system prompts, tool schemas, and environment baggage — ~700 tokens for lean Aider vs ~26,000 for OpenClaw — resent or reconstructed every turn. Teams optimizing only model SKUs are measuring the wrong lever.

Why it matters: Pairs with this week’s Agent Lightning thesis and last week’s AVO “system over model” result without repeating either story: harness design is FinOps and reliability, not just DX.

Source: The New Stack — Agent harness token costs


Aug 17, 2026 – Aug 23, 2026

Weekly AI News Digest

Coverage window: Monday, August 17, 2026 – Sunday, August 23, 2026
Publish date: Monday, August 17, 2026


1. DeepSeek open-sources an agent harness where everything is a plugin

Category: AI engineering / open-source agentic runtimes
Source: The New Stack
Date: August 13, 2026

DeepSeek released DeepSeek Harness as a Node.js developer-preview agent runtime under an MIT license — and the GitHub repo crossed 33,000+ stars within hours. The architectural bet is literal: the model adapter, tool registry, session log, and even the agent loop are replaceable plugins, with “no privileged core to patch.” Composition rides on Cordis, a meta-framework for dynamic dependency wiring described in a Peking University / DeepSeek paper.

Four presets ship out of the box: Standard (filesystem, shell, web search, subagents, plan mode), Minimal (bash + str_replace_editor only), Code (tools exposed as a generated TypeScript SDK so multi-step tool work collapses into one model call), and Creator (Standard plus runtime inspection for authoring custom presets). An append-only session log makes every model-visible input reconstructable — resume, fork, replay, telemetry, and the web UI all hang off that event stream.

Why it matters: Last week’s shortlist was sequence policy (Dogwood) and “every company becomes a tools company.” This week’s AI-engineering lead is the open harness those policies and platforms will plug into — plugin-first, MIT-licensed, and already attracting a community plugin ecosystem.

Source: The New Stack — DeepSeek Harness open-source plugins


2. Per-developer environments were the goal. Agents moved the goalposts.

Category: Platform engineering / AI infrastructure
Source: The New Stack
Date: August 15, 2026

Platform engineering spent the 2020s shrinking the tenant — org → team → developer — until a namespace-per-engineer and seat-based capacity plans looked like the end state. Coding agents broke the math. A developer supervising five parallel agent sessions has five changes in flight; Anthropic’s own C-compiler experiment ran nearly 2,000 Claude Code sessions in two weeks, and Cursor’s docs tell builders to run as many agents as they want in parallel.

The thesis: the tenant is no longer the developer (or even the agent) — it is the change. Seat math underprices change math. A 50-developer org with a few agents per engineer looks like a 300-seat platform on a busy day: each change wants writable data, its own view of shared topics, and a running copy of touched services. Per-developer namespaces, shared staging queues, and headcount-based capacity all misprice that demand.

Why it matters: This is the platform-engineering story of the week opening — IDP tenancy redesigned for agent parallelism, not human headcount. It pairs with last week’s “harness/loop as platform product” thesis and this week’s DeepSeek harness release.

Source: The New Stack — The new tenant is the change


3. Video roundup: agents escaping sandboxes + computer-use agentifies the web + CNCF AI accelerators

Category: AI engineering / platform engineering / cloud-native AI (video)
Sources: YouTube — @PlatformEngineering, @aiDotEngineer, @cncf
Dates: August 10 – August 14, 2026

Primary pick — @PlatformEngineering (August 12): When AI agents escape their sandboxes. Fresh PE-channel framing for containment failure as a platform problem — pairs with the written stack on change-tenancy and plugin harnesses.

Also worth watching this week:

Why it matters: Written coverage this week is open harnesses + change-tenancy; the video circuit is sandbox escape, model→harness governance, and browser/computer-use agents that make those platforms load-bearing.


4. The Human Guide to AI — still waiting on the next chapter

Category: AI literacy / human stories of AI
Source: Medium — The Human Guide to AI
Date: Latest publication post remains July 17, 2026 (no new article since last week)

The publication has not shipped a new piece since The Genesis of Mind (July 17). That article — and the series opener Who is AI? — were already shortlisted in prior digests, so this edition does not rehash either narrative.

This week’s Human Guide slot is a publication watch: the latest available article remains Genesis of Mind. Readers new to the series should start with the publication homepage (Who is AI? → Genesis of Mind). We will feature a full write-up as soon as a newer post lands.

Why it matters: The required Human Guide source is checked without recycling prior shortlists — and the gap itself is a reminder that literacy content moves slower than the harness/model firehose.

Source: Medium — The Human Guide to AI · latest post: The Genesis of Mind


5. Gemini 3.7 Flash hits 65% DeepSWE — and Google cut the Flash price in half (for now)

Category: AI models / enterprise coding agents
Source: The New Stack
Date: August 13, 2026

Three weeks after Gemini 3.6 Flash, Google shipped Gemini 3.7 Flash as the new workhorse for coding agents and longer automated workflows — not the Gemini 3.5 Pro many expected. Google pitches better recovery when agents get stuck and better judgment about when to ask for more information before proceeding.

Benchmarks moved hard: FrontierCode 1.1 Main 34.4% → 43.6%; DeepSWE v1.1 49% → 65.3%. Three thinking_level settings (low / medium / high) trade latency for tool-heavy reasoning. Intro API pricing is $0.75 / $3.75 per million input/output tokens (also applied to 3.6 Flash through Dec 31); both rates double on January 1, 2027.

Why it matters: Same week DeepSeek open-sourced a harness and open-weight labs raced on post-training (GLM-5.3, Qwen3.8-27B, Grok 4.6), Google’s mid-tier Flash line is competing on agent coding reliability + temporary price — the enterprise model story of the week opening.

Source: The New Stack — Gemini 3.7 Flash agents