AI Radar
EN edition
Public · Free

Today's signal is not “one more model got better.” The stack around agents is becoming an operating system for real work

🔭 Main Thread

Today's signal is not “one more model got better.” The stack around agents is becoming an operating system for real work.

At the top of the stack, compute and intelligence keep getting cheaper and more industrial. OpenAI's reported 8GW Ohio campus with NVIDIA/SB Energy makes the constraint explicit: power, financing, and data-center execution. Google DeepMind's Gemini 3.7 Flash pushes coding / web development / knowledge work into a cheaper fast-model tier. DeepSeek V4-Pro, Qwen3.8-27B, and the GPT-5.6 Sol price cut all point in the same direction: usable intelligence keeps sliding down the cost curve.

The more interesting move is in the middle layer. DeepSeek Harness, Cursor Origin, LangSmith Tuned Evaluators, Warp Factories, Vercel Sandbox, AgentCore Payments / x402, and GBrain are all versions of the same thesis: an agent is not a prompt plus a model. It is a workspace, repo sync, memory, evaluation, payments, permissions, sandboxing, and review loop.

The Chinese platform side confirms the demand from a different angle. Bilibili and Rednote are full of “complete beginner / all-in-one” Codex, Claude Code, Cherry Studio, and desktop-agent tutorials. Douyin and Rednote are still rewarding visible AI-made work: short films, prompt templates, learning workflows, and personal productivity stories. That is the crossover today: DeepSeek Harness and Codex are infrastructure stories in the West-facing feed, but they are also mass-market learning-anxiety stories in China.

The practical read: stop teaching “AI tool lists.” Teach how an AI work system is designed, reviewed, improved, and turned into repeatable output.

🎯 Primary

Models and products
- OpenAI — Asana reportedly used Codex to clear a test-system migration estimated at five engineering years in two weeks, for roughly $12k. That is not a demo; it is workflow substitution.
- OpenAI × NVIDIA/SB Energy — the 8GW Ohio AI campus story keeps pushing AI competition into power, capital, and infrastructure operations.
- Google DeepMind — Gemini 3.7 Flash targets coding, web dev, and knowledge work at half the launch price of 3.6 Flash.
- Anthropic — Claude text watermarking and EU AI Act FAQ are a governance signal: frontier labs are wrapping models with traceability obligations.
- DeepSeek — V4-Pro strengthens agent workflows and supports the OpenAI Responses API; DeepSeek Harness ships as an MIT developer preview.
- Qwen — Qwen3.8-27B is a dense multimodal local-workflow play with 262K native context and 1M-context expansion.

Agent infrastructure
- OpenClaw v2026.7.1-2 — a stable release fixing npm plugin metadata compatibility; small but relevant if you think personal agents need boring operational baselines.
- Cline desktop v0.0.14-beta.1 — desktop beta can run alongside stable and previews cloud sessions.
- Apple MLX v0.32.1 — GGUF metadata and build fixes, another small piece in the local-AI economics story.
- Cursor Origin — Cursor is moving into code hosting, repo sync, and agent review. That is a direct move onto GitHub's territory.
- LangSmith Tuned Evaluators — starts with Perceived Error over production traces, claiming 82% lower cost than a frontier-model evaluator.
- Vercel Sandbox — the $1M challenge asks whether agents can escape a Firecracker microVM / host network boundary. Security is now part of the agent product surface.
- Warp Factories — cloud software factories configurable as code, with any model, any harness, and your own evals.
- Perplexity Computer — email becomes an agent surface: forward or cc computer@perplexity.com to start an auditable task.

💰 Investor

  • a16z / Stripe — Will Gaybrick frames agent microtransactions as a new commercial layer of the internet. If agents browse and act for users, tiny payments become infrastructure.
  • a16z / Stripe — internal Stripe examples are harder than the slogan: agents write 30% of code in a week; global tax filing goes out in one-third the time. The response is not less hiring, but more leverage.
  • Sequoia — Continual Learning is the investor version of today's memory/harness theme: higher model IQ is not enough if agent experience resets every turn.
  • Garry Tan — GBrain is positioned as a memory + skills layer for personal agents, connecting Codex, Claude Code, Hermes, OpenClaw, and Postgres/pgvector.
  • Sarah Guo — GitHub stars are becoming easier to distort in an agent-amplified world. Watch dependency graphs, real forks, registry downloads, and private usage instead.

🧠 Sense Makers

  • AINews — the external answer key clusters the day around OpenAI/NVIDIA vertical integration, Stripe–OpenRouter price routing, Cursor Origin, orchestration, and eval harnesses.
  • TLDR — independently lands on Cursor Origin, Anthropic revenue, and deadline dividend scaling. The “agents meet production constraints” read is not just our local frame.
  • SemiAnalysis — reads GPT-5.6 Sol's OpenRouter/Vercel price cut as a distribution and market-share move, not only a pricing move.
  • 机器之心 — a major Chinese AI publication uses HarnessEval to explain why agent performance now depends on tools, environments, memory, and evaluation loops. This is China translating “agent harness” for local builders.
  • 机器之心 — Doubao's “work tasks” feature can operate a Windows virtual desktop; Chinese consumer AI is also moving toward GUI agents for ordinary users.
  • 36氪 — asks what remains scarce when AI can generate everything. For creators, the answer is taste, judgment, and the ability to verify quality.
  • Lenny / Claire Vo — a solo founder uses Codex + ChatGPT to build a fashion brand from sketch to 3D printed gown to ecommerce. This is the solo-company signal in concrete form.

🔨 Practitioner

  • Greg Isenberg — his nine ways to make Claude Code stronger are really an “AI employee work environment” checklist: workspace, memory, brief, ticket, eyes, review, schedule, permissions.
  • Every / Sandcastles — analyzes up to 50 videos per week for topic, format, hook, and script technique, then recommends the next video. This is content-radar-as-product.
  • Every — after token usage rises 230%, the right question is not “how do I cut it?” but “what did I buy, learn, and ship?” That is a good template for solo-founder AI cost reviews.
  • AI Engineer — generated video is cheap enough that the bottleneck moves from generation to taste, distribution, and orchestration.
  • Alex Finn — Qwen3.8-27B on a local RTX 5090 is another sign that local agent-stack economics are getting more interesting.

🔥 Pro Radar

  • Hugging Face trending — DeepSeek-V4-Pro-0813 is a high-heat model, reinforcing today's DeepSeek agent/harness cluster.
  • Hugging Face trending — MiniMax-H3 is high-download image-text-to-video, keeping open-source video generation hot.
  • Juejin — China's developer community is flooded with DeepSeek Harness tutorials: installation, plugin stores, beginner guides, and “95k stars in two days” narratives. This is how infrastructure turns into market education.
  • Product Hunt — OpenTrade applies Claude Code / Codex to an open-source trading harness, a useful sign that coding-agent harnesses are leaking into vertical workflows.
  • Hugging Face Journal Club — Direct On-Policy Distillation points toward cheaper transfer of RL gains, fitting the broader small/local model usefulness trend.

🌶️ Viral Watch

The Chinese viral pool scanned 1124 raw posts and 850 semantic candidates. The strongest patterns were visible AI-made work and “complete beginner / all-in-one” agent tutorials. These posts do not say “AI is powerful.” They show the result first, then hand the viewer a template.

  • Short video · Douyin · AI wuxia fight prompt giveaway — 25.33M likes, platform heat pct 100. The hook is not the prompt; it is the finished rainy bamboo-forest fight scene, with the prompt as the reward.
  • Long video · Bilibili · I used AI to beat the hardest running game — 3.72M views, pct 100. Challenge format plus pain point plus AI reversal gives it a clean story arc.
  • Long tutorial · Bilibili / Rednote · Master Codex in 40 minutes — 1.66M views, with the Rednote version over 100k likes. This is the mass-market side of the Codex / agent workflow signal above.
  • Image / short video · Rednote · AI performs seven human emotions — 97.3k likes, pct 99. It turns an abstract generation capability into a simple emotional test.
  • Long video · Bilibili · Cherry Studio V2 real usage scenarios — 506k views, pct 96. Tool content travels when it is embedded in a real workflow.

This is the full public edition. Want the daily story angles, or a radar built for your market positioning?

See how to subscribe →