Start of day · analyzed 2026-08-02 06:40:53 PT
Morning brief
Sunday, August 2, 2026
Overnight developments and what deserves attention today.
60sources scanned
55new signals
36edge cases kept
8confirmed
ListenEnglish edition
📡 Jin Miao Signals — Morning Brief · 2026-08-02
1. Top 5 — what actually matters today
- Kimi K3 reportedly serves better performance-per-dollar on AMD MI355X than on Nvidia B300 — first credible cross-vendor economics on a frontier open-weights model; if it survives independent replication, the "just buy Blackwell" default becomes a per-workload decision for anyone sizing an inference fleet, and it's the kind of datapoint that gets read against AMD/Nvidia into Monday's open. Self-reported bench — treat as directional, not settled. wafer.ai
- An internal OpenAI model called "Astra" reportedly solved 10 major open math and CS problems — posted by a named OpenAI researcher, not a press release, which is why it matters; the story is no longer "AI does contest math" but "AI closes problems humans left open," and that reframes what a research hire is for. [Rumor] — single social post, no paper, no problem list. @polynoamial (pairs oddly well with the human-only proof landing this week that the Burau representation is faithful at n=4 — arXiv)
- Gemini Robotics ER 2 — embodied reasoning shipped against two real robot bodies (Duo and Apollo) — the embodied-reasoning layer is being productized as a model, not a research demo, which is the step that lets robotics teams stop building their own perception-to-plan stack. Fair warning: this is ONGOING, not overnight — I'm carrying it because a robotics foundation model clears my bar regardless of news cycle, not because it broke last night. Google DeepMind
- Only 8.9% of sites block AI crawlers — and 94.8% are never cited in an AI answer — the open-web bargain quietly inverted: you're paying the training cost and getting none of the distribution. For any founder whose funnel starts with search, that's a channel that's already gone, not going. AI Visibility Index
- Antora closes a $550M Series C for thermal battery storage, explicitly aimed at AI datacenter demand — the marginal dollar in "AI infrastructure" keeps moving downstream from chips to electrons; watch this as the tell for where 2027 compute actually gets sited, and as context for power/industrial names levered to datacenter buildout. [Rumor / ONGOING] — round size not primary-sourced. Crunchbase News
2. New-direction sparks
- Frontier weights are collapsing into consumer memory faster than anyone's roadmap assumed. Overnight on r/LocalLLaMA: llama.cpp landed MTP/DSpark support for DeepSeek V4-Flash, a claimed 284B V4-Flash run in 5.3 GB, Kimi K3 pushed onto a single CPU with 8 GB (we were at 29 GB yesterday morning), and 12–15 tok/s on 3090s and MI50s. Non-obvious because none of it is a model release — it's runtime plumbing, the least-covered and most leverage-dense layer. If this holds, the unit of "who can serve a frontier model" moves from a cluster to a laptop inside a quarter. All [Rumor], all self-reported, no links in feed — r/LocalLLaMA threads.
- "The Greenhouse and the Lens" — a working taxonomy for two distinct modes of agentic work. Rare thing: someone naming the shape of agent collaboration instead of shipping another harness. Non-obvious because the field is drowning in tooling and starved of vocabulary, and vocabulary is what lets teams argue productively about which mode they're in. brethorsting.com
3. Threads worth watching
- Human–AI interaction — moved materially, from an unlikely source. Greg Brockman: at OpenAI, people hook ChatGPT into Slack, and colleagues hate it — they'd happily do the same task if a human asked. The rejection isn't about capability, it's about relationship. That's the sharpest empirical read I've seen on the ceiling of agent-to-human delegation, and it came out as a throwaway quote. simonwillison.net
4. Contrarian watch
- Nvidia's inference moat is being priced as permanent; the MI355X number says it's per-workload. One blog post isn't a thesis, but it's the first cross-vendor perf/$ claim specific enough to be falsified. Context only, but note the Guardian's read this morning that market turmoil is exposing how opaque the AI economy's actual unit costs are. Guardian
- Consensus: frontier capability requires datacenter capital. Edge: 284B params in 5.3 GB. If even half the local-inference claims replicate, the "compute is the moat" argument is weaker at the serving layer than the capex narrative implies. Watch whether these get reproduced by anyone with a name attached this week.
- The EU AI Act's general applicability date is today, 2026-08-02. Circulating on r/LocalLLaMA as a punchline; it is not a punchline for anyone shipping into the EU. No link in feed — verify against the official Official Journal text before you act on any of it, including mine.
- AI firms are reportedly scanning rare book editions and destroying the physical copies afterward. Non-consensus because the training-data fight has been framed entirely as copying; this frames it as loss. Different legal theory, different constituency, much worse optics. Dallas Express
5. Verification flags
- ⚠️ do not act on yet — needs primary source: OpenAI "Astra" solving 10 major open math/CS problems. Single social post, no problem list, no paper. @polynoamial
- ⚠️ do not act on yet — needs primary source: Antora's $550M Series C size and lead investor — secondary reporting only. Crunchbase News
- ⚠️ do not act on yet — needs primary source: every local-inference number above (V4-Flash in 5.3 GB, K3 on 8 GB CPU, 12–15 tok/s figures) — anonymous, self-reported, no reproducible configs.
- ⚠️ do not act on yet — needs primary source: MI355X-vs-B300 perf/$ — vendor-adjacent blog, no third-party replication. wafer.ai
- ⚠️ do not act on yet — needs primary source: EU AI Act applicability details as described in social chatter — read the statute, not the thread.
Markets context only — not financial advice.
Listen中文音频
📡 Jin Miao Signals — 晨间简报 · 2026-08-02
1. 今日五条真正重要的消息
- 有报告称 Kimi K3 在 AMD MI355X 上的"每美元性能"优于 Nvidia B300 —— 这是首份针对前沿开源权重模型、可信度尚可的跨厂商经济性数据。如果它能扛住第三方复现,那么"闭眼买 Blackwell"这个默认选项,对任何要规划推理集群的人来说都将变成一道逐场景计算的选择题;这类数据点也很容易在周一开盘时被拿来对照 AMD 与 Nvidia。注意:厂商自测跑分,只作方向性参考,远未定论。wafer.ai
- OpenAI 内部模型 "Astra" 据称攻克了十个数学与计算机领域的重大公开难题 —— 消息来自一位具名 OpenAI 研究员的社交发帖,而非官方通稿,这恰恰是它值得关注的原因:故事线不再是"AI 会做竞赛数学题",而是"AI 解决了人类留下的开放问题",这将重新定义"招一名研究员"到底意味着什么。[传闻] —— 仅有一条社交发帖,无论文,无问题清单。@polynoamial (与本周落地的另一项纯人类证明形成微妙呼应:Burau 表示在 n=4 时是忠实的 —— arXiv)
- Gemini Robotics ER 2 —— 具身推理能力在两款真实机器人本体(Duo 与 Apollo)上落地 —— 具身推理层正在被当作一个模型来产品化,而不是一个研究演示。这一步意味着机器人团队可以不必再自建"感知到规划"的整套技术栈。先说清楚:这是持续追踪项,不是昨夜突发。我把它放进来,是因为机器人基础模型本身就够格上榜,与新闻周期无关。Google DeepMind
- 只有 8.9% 的网站在屏蔽 AI 爬虫 —— 而 94.8% 的网站从未被 AI 答案引用过 —— 开放网络的那笔交易已经悄然反转:你付出了训练语料的成本,却拿不到任何分发。对任何漏斗起点是搜索的创业者来说,这个渠道不是"正在消失",而是"已经没了"。AI Visibility Index
- Antora 完成 5.5 亿美元 C 轮融资,热储能业务明确瞄准 AI 数据中心需求 —— "AI 基础设施"里的边际资金正持续从芯片向下游的电子流动。把它当作 2027 年算力究竟落在哪里的风向标,也当作理解那些绑定数据中心建设的电力与工业标的的背景。[传闻 / 持续追踪] —— 融资规模未获一手信源证实。Crunchbase News
2. 新方向的火花
- 前沿模型权重正在以超出所有人路线图预期的速度,塌缩进消费级显存里。 昨夜 r/LocalLLaMA 上:llama.cpp 合入了针对 DeepSeek V4-Flash 的 MTP/DSpark 支持;有人声称 2840 亿参数的 V4-Flash 只用 5.3 GB 就跑了起来;Kimi K3 被压进单颗 CPU、8 GB 内存(昨天早上这个数字还是 29 GB);3090 和 MI50 上跑出 12–15 tok/s。之所以不那么显眼,是因为这里没有一项是模型发布 —— 全是运行时的管道工程,报道最少、杠杆最高的那一层。如果这些说法站得住,"谁能提供前沿模型服务"的门槛单位,将在一个季度内从集群降到一台笔记本。以上全部为[传闻],全部自测自报,信息流中无链接 —— 来源为 r/LocalLLaMA 帖子。
- 《温室与透镜》—— 一套区分两种智能体工作模式的实用分类法。 难得:终于有人去命名智能体协作的形态,而不是再造一个框架。它不显眼,是因为这个领域工具泛滥、词汇匮乏,而恰恰是词汇,才能让团队就"我们现在处于哪种模式"展开有效争论。brethorsting.com
3. 值得追踪的线索
- 人机交互 —— 有实质进展,且来源出人意料。 Greg Brockman 说:在 OpenAI 内部,有人把 ChatGPT 接进了 Slack,结果同事们很反感 —— 而同样一件事,如果是真人开口请他们做,他们乐意得很。这种抗拒无关能力,关乎关系。这是我见过对"智能体向人类委派任务"天花板最锐利的一次经验性判断,而它只是随口一句话。simonwillison.net
4. 逆共识观察
- 市场把 Nvidia 的推理护城河定价为永久性的;MI355X 那组数字说,它其实是逐场景的。 一篇博客构不成论点,但这是第一份具体到足以被证伪的跨厂商性价比声明。仅作背景参考,但值得留意《卫报》今早的判断:市场动荡正在暴露 AI 经济的真实单位成本有多不透明。Guardian
- 共识:前沿能力需要数据中心级资本。边缘信号:2840 亿参数,5.3 GB。 哪怕本地推理的说法只有一半能复现,"算力即护城河"在服务这一层的成立程度,也远低于资本开支叙事所暗示的。本周留意是否有具名者能把这些结果复现出来。
- 欧盟《AI 法案》的普遍适用日就是今天,2026-08-02。 它在 r/LocalLLaMA 上被当作一句段子传播;但对任何要把产品卖进欧盟的人来说,它不是段子。信息流中无链接 —— 在你据此采取任何行动之前,请以官方公报原文为准,包括我这条也一样。
- 有报道称,AI 公司在扫描珍本古籍后销毁了纸质原件。 之所以是非共识,是因为训练数据之争此前完全被框定为复制;而这件事把它框成了灭失。不同的法律逻辑,不同的利益相关方,观感也差得多。Dallas Express
5. 待核验清单
- ⚠️ 暂勿据此行动 —— 需一手信源:OpenAI "Astra" 攻克十个数学/计算机重大公开难题。仅一条社交发帖,无问题清单,无论文。@polynoamial
- ⚠️ 暂勿据此行动 —— 需一手信源:Antora 5.5 亿美元 C 轮的融资规模与领投方 —— 仅有二手报道。Crunchbase News
- ⚠️ 暂勿据此行动 —— 需一手信源:上文全部本地推理数字(V4-Flash 跑在 5.3 GB、K3 跑在 8 GB CPU、12–15 tok/s)—— 匿名自报,无可复现配置。
- ⚠️ 暂勿据此行动 —— 需一手信源:MI355X 对比 B300 的性价比 —— 出自与厂商关系密切的博客,无第三方复现。wafer.ai
- ⚠️ 暂勿据此行动 —— 需一手信源:社交讨论中所描述的欧盟《AI 法案》适用细节 —— 请读法条,别读帖子。
仅为市场背景信息,不构成投资建议。
Private founder layer
Co-founder confidential
Strategic synthesis and adversarial review, encrypted in the page source.
That passphrase did not decrypt this edition.
Confidential · English
机密内容 · 中文
Source ledgerEvery scored item, including outliers
- i5 / e5
- I pushed Kimi K3 onto one CPU with 8 GB of RAMreddit/r/LocalLLaMAi5 / e5
- DeepSeek-V4-Flash 284B on 5.3GB of memoryreddit/r/LocalLLaMAi5 / e5
- llama.cpp just added MTP / DSpark support for DeepSeek V4 Flashreddit/r/LocalLLaMAi5 / e5
- Setting up of a 16xGB10 (DGX Spark) clusterreddit/r/LocalLLaMAi4 / e5
- PSA for DeepSeek-V4-Flash-0731 users — don't blow out your prompt cache with system role messages mid-conversationreddit/r/LocalLLaMAi4 / e5
- Ran DS V4-Flash-0731 Locally on 3xMI50 32GB @ ~15 t/s TGreddit/r/LocalLLaMAi4 / e5
- i5 / e4
- i4 / e4
- i4 / e4
- i4 / e4
- DeepSeek-V4-Flash-0731 UD-IQ3_S 12.5 tok/s on RTX 3090 +128GB DDR5reddit/r/LocalLLaMAi4 / e4
- i4 / e4
- i4 / e4
- i4 / e4
- i2 / e5
- i3 / e4
- Agent4Leasehackernewsi3 / e4
- Vacuum 16Treddit/r/LocalLLaMAi3 / e4
- Xberg v1 is outreddit/r/LocalLLaMAi3 / e4
- i3 / e4
- i3 / e4
- i4 / e3
- EU AI Act takes effect tomorrow, August 2, 2026. 🤡reddit/r/LocalLLaMAi4 / e3
- i2 / e4
- i2 / e4
- i3 / e3
- TimeOS 2.0rssi3 / e3
- Anime User Interfaceshackernewsi2 / e3
- No replies to rebuttals and comments even by AC [D]reddit/r/MachineLearningi2 / e3
- Capptivorssi2 / e3
- i2 / e3
- Zen Whisperrssi2 / e3
- Termexorssi2 / e3
- Bolcho AIrssi2 / e3
- Finamierssi2 / e3
- i3 / e3
- i3 / e3
- i2 / e3
- The Art of 64-bit Assemblyhackernewsi2 / e3
- i2 / e3
- i2 / e3
- i3 / e2
- Diátaxishackernewsi3 / e2
- A big win for Android interoperabilityhackernewsi3 / e2
- I don't recommend Tailwind CSShackernewsi2 / e2
- Show HN: Elevatorshackernewsi2 / e2
- i2 / e2
- Seedance 2.5hackernewsi2 / e2
- NetBSD 11.0hackernewsi2 / e2
- Go 1.27 Interactive Tourhackernewsi2 / e2
- i2 / e2
- i2 / e2
- i2 / e2
- i1 / e2
- ARR August Cycle [D]reddit/r/MachineLearningi1 / e2
- Question about NeurIPS discussion phase [D]reddit/r/MachineLearningi1 / e2
- Discussion About Meta-Review Issue Report in ARR Cycle [D]reddit/r/MachineLearningi1 / e2
- i1 / e1
- i1 / e1