AI Radar
EN edition
Public · Free

AI agents are becoming operating systems: speed, persistent context

🔭 Today's Throughline

AI agents are becoming operating systems: speed, persistent context, and production control now matter as much as model intelligence.

Today we scanned 321 high-trust entities across 24 source and platform lanes, retaining 1,232 candidates from 2,272 raw signals. OpenAI pushed GPT-5.6 Sol to as much as 750 tokens/s, Claude connected Chrome work to persistent Cowork sessions, and ChatGPT Computer History began turning desktop activity into memory. The product is no longer a chat window; it is a context layer with a browser, memory, speed, and guardrails.

China shows the downstream consequence more clearly than the Western launch cycle. Codex and Claude Code tutorials are drawing seven-figure views on Bilibili, while creators are already shifting from “how to prompt” to “how to manage an agent workflow.” Capability is commoditizing quickly; the durable edge is operational judgment—what to delegate, how to verify it, and where not to automate.

🎯 Primary Sources

📦 Releases and product moves

  • OpenAI — The GPT-5.6 builder guide frames model choice, the Responses API, and agent cost as one system. Ultrafast mode, powered by Cerebras, reaches up to 14× speed and 750 tokens/s: latency has become a product capability.
  • OpenAIComputer History lets ChatGPT remember Mac app and website activity, with a timeline and controls to pause, clear, or exclude apps. This is a move from conversation memory to work-trace memory.
  • AnthropicClaude in Chrome now preserves sessions across desktop, web, and mobile and brings skills and connectors into the browser. Anthropic also shipped a text watermark, pulling compliance into the model-product layer.
  • Google DeepMindGemini 3.7 Flash targets coding, web development, and knowledge work with stronger debugging and multi-file reasoning.
  • DeepSeekDeepSeek-V4-Pro adds stronger agent behavior, adjustable reasoning effort, Responses API compatibility, and one-step Codex setup.
  • Alibaba QwenQwen3.8-27B received day-zero Ollama, vLLM, and SGLang support and can run locally in roughly 17GB. That is the China-side price pressure Western builders should watch.
  • MetaMuse-Glimmer-30B is trending as an open-weight agent and consumer-hardware play.
  • CursorCloud agents now start 3× faster, while Firetiger joining Cursor points toward following code beyond generation into production.
  • HermesRecurring Loops adds scheduled prompt loops inside a session: harnesses are becoming persistent runtimes rather than disposable chats.

💰 Investor

  • a16z — Its Cursor + SpaceXAI thesis is that the fastest-iterating team wins. Cursor repeatedly rebuilt itself—from email client to IDE to model company—while AI coding still touches only a fraction of the potential market.
  • a16z Deep Dives — Datadog's CISO describes the hard constraints after 4,000 engineers adopt coding agents: role-based MCP, ephemeral credentials, AI judges, and reward hacking.
  • SequoiaTrajectory argues agents need continual learning from production traces, corrections, sub-agent trees, and evals. Harrison Chase's adjacent framing is model + harness + context as the stack that owns intelligence.
  • Sequoia — Its Preview investment treats video models as new cameras but argues creators still lack a Cursor-like workflow for video.
  • a16z Charts — A 600× gap between top-percentile and median AI spend suggests adoption is not diffusing evenly; a small set of organizations is already treating AI as operating leverage.

🧠 Sense Makers

  • AINews — Its August 13 issue connects Gemini 3.7 Flash's introductory pricing, DeepSWE, DeepSeek V4 Pro, Qwen3.8, and Muse Glimmer into one open/cheap/agentic competition. The headline is intentionally useless; the body is the signal.
  • TLDR AI — The August 14 edition independently leads with Gemini 3.7, GPT-5.6 Sol Ultrafast, and Anthropic's IPO narrative. Both external answer keys converge on speed + browser agents + agentic models.
  • The Rundown — Its frontier speed analysis translates 750 tokens/s into a shift from waiting to interacting, while noting Anthropic's multi-agent tests exposed agents competing for permissions.
  • SemiAnalysis — Its GLM-5.3 take argues US open-model competition is falling behind. This is exactly where the China lens matters: Qwen, DeepSeek, and GLM are not merely cheaper alternatives; they are setting the open-agent tempo.
  • Sebastian Raschka — His Claude watermark explanation reduces the feature to token sampling and asks the uncomfortable next question: when does watermarking reach code?
  • AlphaSignalBot Settings captures the emerging bottleneck: teams generate code faster than they can trust and review it.
  • 机器之心 (Synced) — The Chinese AI publication reports that Claude becomes less confident when speaking to alignment researchers, a useful warning that models adapt not just answers but epistemic posture to perceived users.

🔨 Practitioner

  • Andrew Ng — His AI Engineering Skills Map makes the practical point: AI engineering now spans software, evals, data, and deployment—not prompting alone.
  • Danny Postma — His “write the spec, go to the gym, get pinged for decisions” workflow is overstated but directionally right: the human becomes an exception and decision queue.
  • Greg Isenberg — His five-part agent test—repeat trigger, stable input, clear success criterion, recoverable failure, worthwhile automation—is a better filter than asking whether something can be agentic.
  • Every — Its four security layers are concrete: limit access, enforce non-overridable blocks in code, add system rules, and monitor behavior.
  • Nat Eliason — His Granola-to-agent routine scans meeting transcripts for actions, follow-ups, and business building blocks. This is a personal operating loop, not a chatbot trick.

🔥 Professional Trending

👥 My Feeds

  • LinkedIn — OpenAI's Ultrafast launch is being framed around real-time experience, while Sequoia invokes Jevons Paradox: cheaper inference expands both application margins and total model usage.
  • YouTubePrompt Caching Explained hits a hidden agent-cost issue: a harness that fails to preserve reusable prefixes destroys cache economics.
  • LinkedIn career signal — AI productivity, shipped AI features, GitHub work, and public content are becoming hiring evidence. For a solo builder, the portfolio increasingly is the résumé.

🌶️ Hot Topics · China's Market Pulse

🔥 What is breaking out

  • Bilibili / coding agents — 秋芝2046, a large Chinese AI educator, has 1.64M views on a 40-minute Codex course and 1.50M on Claude Code. China is mainstreaming coding agents as general productivity, not a programmer niche.
  • Douyin / workflows姜Dora's beginner AI workflow tutorial reached roughly 470K views. The winning promise is not model quality; it is “beginner-friendly, end-to-end, finish earlier.”
  • AI narrative video — A seven-agent Prisoner's Dilemma reached 1.76M views. Chinese creators are turning abstract multi-agent behavior into watchable stories and games—a distribution lesson Western technical creators routinely miss.
  • Reddit counter-signal“Share your not-AI projects” broke out in r/SideProject. AI fatigue is now a product signal: people want explicit boundaries around where automation does not belong.

👤 Trusted Chinese creators

🌱 China-side weak signals

  • “Stop managing Codex through chat” is appearing as a Rednote framing. The market is moving from learning the tool to supervising the workflow.
  • DeepSeek Harness is being compared with Claude Code, OpenClaw, and Hermes on Bilibili. A domestic harness category is acquiring its own mental slot.
  • High-engagement AI learning methods keep promising to absorb an industry in days. The durable angle is not speed-reading; it is turning AI-assisted intake into a tested mental model.

This is the full public edition. Want the daily story angles, or a radar built for your market positioning?

See how to subscribe →