W20

Local AI 浪潮、LLM 自毀文件、Claude Opus 4.7 GA — Big Tech 信任度繼續鬆動

W20 兩條敘事疊在一起:一邊是 Anthropic 把 Opus 4.7 推 GA(軟體工程顯著躍進、長任務 reliability、視覺解析度升級),Vercel 把 Opus 4.7 fast mode + AI Gateway production index 推上線;另一邊 Hacker News 一週連發三條 Big Tech 不信任訊號 — Google 把 reCAPTCHA 對 de-Google Android 用戶弄壞、Meta 關 IG E2E 加密、Chrome 偷裝 4GB 模型尾音持續Hacker News 同週開始討論「Local AI 該是常態」、「LLM 委派文件會悄悄改壞檔案」、「ChatGPT 5.5 Pro 對 Fields Medalist 數學的限制」、Louis Rossmann 出錢替 OrcaSlicer dev 打官司a16z 連發 4 篇思路文章(System of Record → System of Intelligence、Stitch、Software Losing Its Head、American Technology Ships Values)

297篇文章
13+來源
Y

AI 模型 & 產品更新

Claude Opus 4.7 — 從 5/9 release 走到 5/11 GA

Anthropic W19 推 Opus 4.7,W20 進一步走完 GA + 商業化動作重點不是「又一個版本」,是幾條 reliability 訊號疊在一起:

  • 軟體工程 benchmark 對 4.6 顯著升級,最難的任務拉得最多
  • Hex / Notion / Replit / Vercel / Genspark / Warp 早期測試者共識:長任務不會中途垮掉(loop resistance、tool error 變少、能撐幾小時不放棄)
  • Vision 解析度拉到 2576px(4.6 是 1152px)— 圖表、技術 schematic、Figma 截圖可以直接餵原圖
  • Cyber 能力刻意降級 + safeguard:Anthropic 第一次明說「先在較弱模型上實驗 cyber 阻擋機制,再放更強的 Mythos」

我的觀察: 對接案工作者來說,重點不在「4.7 比 4.6 多 X%」,是「hand off 信心」這個敘事被驗證了 — 早期測試者形容「以前要盯著做的事,現在敢交給 Opus 跑」如果你接案模式是把 AI 當執行夥伴,這個訊號意味著「可委派任務」的 surface area 拉大了一輪

Anthropic 5/14:Claude for Small Business(HN 上線)

Anthropic 5/14 把 Claude 包成「Small Business」product line — 不只是 ChatGPT Team 的對打版,是把使用情境往「小公司/freelance/顧問」這個 segment 直接定位對 AI 公司產品線的意義:以前是 dev tool(API + Claude Code)+ enterprise,現在把中段加進來

Meta 5/16 三條同日:MTIA 第二代、SAM 3.1、Muse Spark

Meta AI blog 5/16 一次上四條主要內容:

  • MTIA 兩年四晶片:自研晶片路線 confirmed,scale to billions
  • SAM 3.1:multiplexing 讓 16 物件一次 forward pass,影片處理速度翻倍(16 → 32 fps on H100),real-time 物件追蹤門檻降低
  • Muse Spark:MSL 推「personal superintelligence」概念,contemplating mode(多 agent 平行推理)對打 Gemini Deep Think / GPT Pro,Humanity's Last Exam 58% / FrontierScience Research 38%
  • Scaling 框架白皮書:把 frontier risk evaluation / loss of control 寫成 published framework,是這波 Anthropic + Meta 同時往「safety transparency」靠的訊號

我的觀察: Muse Spark 對 personal AI 這條敘事繼續加碼「personal」是 framing — 不是 model 本身,是 model 服務的對象對自建 wiki/個人 ops 系統這條路是利多

Anthropic 5/14 ↔ Microsoft Foundry / Bedrock / Vertex AI

Opus 4.7 同日對 AWS Bedrock、GCP Vertex AI、Microsoft Foundry 全平台同步上線價格不變($5 / $25 per million)意思是 enterprise 採購不會卡單一供應商,rate limit pressure 分散到三家 hyperscaler


AI 開發工具 & Agent

Claude Code 一週連發 v2.1.136 → v2.1.142(5/9 → 5/15)

Claude Code 倉庫一週 7 個 release,幾個值得記的:

版本重點
v2.1.142claude agents 加 flags:--add-dir、--mcp-config、--plugin-dir、--permission-mode、--model、--effort、--dangerously-skip-permissions;fast mode 預設 Opus 4.7
v2.1.141Hook 加 terminalSequence field(桌面通知 / window title / bell 不需要 controlling terminal);CLAUDE_CODE_PLUGIN_PREFER_HTTPS 用 HTTPS clone plugin 而非 SSH
v2.1.140subagent_type 大小寫不敏感("Code Reviewer" → code-reviewer);/goal hang fix
v2.1.139Agent View — 一個面板看所有 session(running / blocked on you / done);新 /goal 指令(設完成條件後 Claude 跨 turn 持續做)
v2.1.136autoMode.hard_deny — auto mode 的硬擋規則(無視 intent / allow exception)

我的觀察: v2.1.139 的 Agent View 是這週對「並行使用 AI」最大的工程進步我自己跑 background agent 寫補助案 / sync events / 跑 reflect cron,過去要靠 tmux 加 git log 拼湊,現在是 first-class UI/goal 是另一條 — 對 long-horizon task(30 分鐘以上)的可預期性是分水嶺

Vercel:Opus 4.7 Fast Mode + AI Gateway Production Index(5/12)

Vercel 5/12 一次推三件:

  1. Opus 4.7 Fast Mode(research preview)— output token 生成速度 ~2.5x 快,intelligence 不變speed: 'fast' 在 anthropic provider options 切換
  2. AI Gateway Production Index — 用 Vercel 自家 production workload 排序當週哪個 model 最強意思不是 benchmark,是「真實流量上誰好用」
  3. Vercel Firewall 用自然語言寫 WAF rule(CLI / dashboard 都有)

我的觀察: Fast Mode 對接案場景很實際 — Opus 4.7 跑 sales page / proposal / 簡報草稿時,2.5x 速度差別等於整個下午的時間槓桿AI Gateway production index 比 benchmark 可靠,因為它 reflect 真實 mix 工作量

Vercel Trusted Sources for Deployment Protection(5/13)

Vercel 推 OIDC-based「Trusted Sources」取代 long-lived Protection Bypass secret對 CI/CD 接 Vercel preview 的安全性是大升級 — 不再需要在 GitHub Actions 存長期 token

Vercel Sandbox:Node.js 26 + 自訂 proxy(5/11 + 5/12)

Sandbox 加 Node 26 支援、firewall 支援 request proxying / filtering對需要在隔離環境跑用戶 code 的場景(教育平台 / 程式練習 sandbox / agentic IDE)門檻降低

Hacker News:Local AI 該是常態(5/11)

HN 一篇「Local AI needs to be the norm」討論度很高論點:On-device model 不該是奢侈品,是隱私 default對應 Chrome 偷裝 4GB 模型的爭議 — 重點不是「不裝」,是「讓使用者知道並選擇」

我的觀察: 接案場景對 local AI 的需求其實比 SaaS 高 — 客戶資料、會議錄音、合約草稿,能在本機跑就少一個外洩 vectorApple Foundation Models / Ollama / LM Studio 這條路 W20 比 W19 多了一波 mainstream attention

Hacker News:LLM 委派文件會悄悄改壞檔案(5/9)

「LLMs corrupt your documents when you delegate」討論意思是 LLM 改 docx / pdf / 結構化文件時,會把 metadata / 公式 / 格式悄悄改壞,使用者打開看內容沒問題,過幾週才發現某個欄位數字跑掉

我的觀察: 對接案者最直接的影響是報價單 / 合約 / 提案委派 LLM 改文件後要走「diff 對比」+「人工最後核對關鍵數字」這個 checklist,不能裸信「LLM 說改完了」

Hacker News:ChatGPT 5.5 Pro 對 Fields Medalist 的限制(5/9)

W19 已報過 Gowers 親評W20 後續討論在 HN 延燒 — 重點不是模型解不解得了,是研究級數學家持續 publicly 寫 LLM 心得這件事本身

Hacker News:Claude Code 寫 HTML 的「不合理有效」(5/9)

W19 帶過W20 在 HN 主流 thread 繼續發酵 — 接案 / sales page / proposal page 越來越多人放棄 React/Vue 抽象,直接讓 Claude Code 寫 raw HTML


大神觀點

Anthropic
Anthropic — Opus 4.7 GA + Claude for Small Business + Cyber Verification Program — 把 SMB 段加進產品線,並第一次明說「弱模型先實驗 cyber safeguard,再放 Mythos」
Boris Cherny / Claude Code
Boris Cherny / Claude Code — 一週 7 個 release(v2.1.136 → v2.1.142)— Agent View 是並行 session 的 first-class UI,/goal 把 long-horizon task 變可預期
Meta AI / MSL
Meta AI / MSL — 同日推 MTIA 第二代 + SAM 3.1(影片處理速度翻倍)+ Muse Spark contemplating mode + Advanced AI Scaling Framework — 把 safety transparency 寫成 published framework
Andrew Ng (The Batch)
Andrew Ng (The Batch) — 過去兩個月 weekly digest 連發 10 篇補完(Mar 13 → May 15)— 從 GPT-5.4 splash 到 Claude Code source leak 到 GPT-5.5 Outperforms(and Hallucinates)一次補課
a16z / Andreessen
a16z / Andreessen — 5/12-5/15 連發 4 篇:System of Record → System of Intelligence、Stitch 投資、Software Losing Its Head、American Technology Ships Values — 把「軟體形態正在轉換」做成連續敘事
Dan Shipper (Every)
Dan Shipper (Every) — 「Socrates as a Service」— 跟 Drew Bent Learning Mode、Anthropic Teaching Claude Why 同一波「不直接給答案,幫使用者思考」的思潮
Louis Rossmann
Louis Rossmann — 5/10 公開出錢替 OrcaSlicer 開發者打官司(被 3D 列印公司威脅)— 開源維護者法律風險變主流話題

創投 & 市場

a16z 連發四篇(5/12 → 5/15)

a16z 一週四篇連發,可以串成一條敘事:

  1. 5/12「No Man Left Behind」: American Technology Ships with Our Values — 美國科技帶價值觀出口,sovereign AI / national AI 框架繼續加碼
  2. 5/14 Is Software Losing Its Head? — 軟體形態正在從「app + UI」轉成「agent + intent」
  3. 5/14 Investing in Stitch — Andreessen Horowitz 投 Stitch(agentic infra 賽道)
  4. 5/15 From "System of Record" to "System of Intelligence" — Workday / Salesforce / SAP 這代 SaaS 是 system of record,下一代是 system of intelligence

我的觀察: 這四篇串起來是「SaaS 形態 reboot」的論點 — 過去 20 年的 SaaS 護城河(record + workflow)正在被 agent + intent 重寫對接案者意義:客戶現在問的 chatbot / AI assistant / agent 不是 nice-to-have add-on,是下一代 software 的 entry point

First Round Review 同週 10 篇 PMF 集中發

First Round 一週連發 10 篇文章,其中幾條對接案者特別有意思:

  • "AI-Powered" Isn't a Position — 把「AI-powered」當賣點等於沒定位
  • Forward Deployed Engineer — Palantir-style 駐點工程師角色,技術 + business + 關係三線
  • Discovery Toolkit — 研究思維加速產品探索期
  • Reluctantly Influential: Lenny Rachitsky — 「不情願的影響力」典型 freelance / solopreneur 路徑

我的觀察: "AI-Powered" Isn't a Position 是這週最值得反覆讀的一篇 — 對任何在做提案 / case study / pitch 的人,把「AI」當賣點本身就是放棄定位客戶買的不是 AI,是某個問題被解決

Reuters / Finimize:總體面繼續震盪

W20 Reuters / Finimize 共 196 篇主要圍繞:Fed Warsh 接任後 rate hike 預期、油價測 $110、美元走強、UK 政治壓力 + 油價推升通膨對接案者不是直接訊號,但 client 預算心理(特別是企業客戶)會持續保守 — 提案要往「省錢 / 自動化 / ROI 明確」三個方向校準

Project Zero:Pixel 10 0-click exploit chain(5/15)

Google Project Zero 揭露 Pixel 10 的 0-click exploit chainHN 125 分對手機作為「主權設備」這個假設是反向訊號 — 即使最 secure 的 Android 也會被 break

Antón Leicht:「Access to frontier AI will soon be limited by economic and security constraints」(HN 194↑)

論點:前沿 AI 不會永遠 democratized — 因為算力 / 能源 / 安全顧慮,最強的 model 會變稀有資源對「人人都用 GPT-5 / Opus 4.7」這個 default assumption 是反向訊號

UK Sovereign LLM Inference: relax.ai(HN 98↑)

英國 sovereign LLM inference 出現獨立 player台灣 sovereign AI 提案敘事可以引這條 — 不只主權國家在做,民間獨立公司也在做


行動建議

  1. Opus 4.7 hand-off 信心測試(30 min)— 找一個你過去不敢交給 AI 跑的長任務(補助案章節重寫 / 全 repo refactor / 一週 events.jsonl 跨日推理),全程不干預 Opus 4.7 跑完看結果決定下一個月哪些任務移到「可委派」清單
  2. Claude Code Agent View 試用 + /goal 設定(20 min)— v2.1.139+ 升級後,把現有 background agent(reflect / dream / sync)改用 Agent View 管,long-horizon task 用 /goal 設完成條件
  3. Vercel Opus 4.7 Fast Mode 試裝(15 min)— 拿一個既有的 AI Gateway 接案專案切 speed: 'fast',量化體感差異如果 2.5x 是真的,下次接案默認開
  4. 「AI-Powered」反 positioning 自檢(30 min)— 把過去三個月的 proposal / pitch deck / personal site 文案 grep「AI-powered」「AI-driven」「AI-enabled」這類句子每一句改成具體問題 statement
  5. 本機 AI workflow pilot(1 hr)— 把「客戶會議錄音 → transcript → 摘要」這條 pipeline 從雲端模型切到本機(Whisper.cpp + 本機 Llama / Ollama)對客戶資料外洩 vector 是 immediate 削減
  6. LLM 委派文件加 diff check(10 min)— 提案 / 合約 / 報價單委派 LLM 改後,加一條 SOP:用 git diff 或 docx diff 工具看完整改動,不裸信 LLM 報告
  7. 「System of Record → System of Intelligence」敘事補進提案(30 min)— a16z 5/15 這條論點對企業客戶提案是好 framing客戶問「為什麼要加 AI」時,給的不是 feature list,是「你現在的系統是 record,AI 變成 intelligence layer」
  8. Forward Deployed Engineer 角色補進對外文案(10 min)— 比「全端工程師」「AI 顧問」精準下次接案 / 找合作 / 寫 LinkedIn bio framing 直接套

參考來源

RSS Digest: 297 articles from 13 sources(W20,5/10-5/16,7 days)

主要訊號分布:

  • Economy & Finance: 196 篇(Reuters Business / Finimize 為主)
  • Cloud Infrastructure: 37 篇(Google Cloud + Vercel + Azure + Meta AI infra)
  • Hacker News Top: 23 篇(Local AI / LLM doc corruption / Pixel 10 exploit / UK sovereign)
  • AI Engineering: 17 篇(The Batch 10 篇補課 + Claude Code 7 release)
  • Business & Startups: 10 篇(First Round Review batch)
  • AI Companies: 8 篇(Anthropic Opus 4.7 GA + Meta 五連發)
  • Builders & Indie Hackers: 6 篇(a16z 4 篇 + Dan Shipper 2 篇)

今日 anchor 文章:

Generated from W20 digest(2026-05-16 00:12)+ 補章節 2026-05-17

Opus 4.7Local AIBig Tech TrustClaude CodeVercelHacker Newsa16zPrivacy