In this issue:
GitHub Squad: Multi-Agent Inside Your Repo Cursor Built on Moonshot AI's Kimi Flash-MoE: 397B Model on a Laptop OpenAI Building Automated Researcher Dreamer: Personal Agent OS Simon Willison: Git + Coding Agents LangChain Fleet + Sandboxes Iran War Threatens AI Chip Supply Chain Action Items Sources
GitHub Squad: Multi-Agent Inside Your Repo
GitHub released Squad — a system for running multiple AI agents collaboratively inside a repository, extending GitHub Copilot
Unlike existing multi-agent frameworks, Squad emphasizes three design principles:
| Principle | Meaning |
|---|---|
| Inspectable | Every agent decision can be reviewed by humans |
| Predictable | Agent behavior follows predictable patterns |
| Collaborative | Agents work together, not in isolation |
This contrasts with last week's Pragmatic Engineer report that "AI agents are slowing us down." GitHub's response: the problem isn't agents themselves, but the quality of orchestration
What this means for us: Our Claude Code system already runs 15+ agents, but orchestration remains the biggest challenge. Squad's "inspectable" principle is worth adopting — every subagent decision should be traceable
Cursor Admits New Model Built on Moonshot AI's Kimi
Cursor launched Composer 2, but it was quickly discovered that the underlying model is based on Chinese company Moonshot AI's Kimi-k2.5
TechCrunch's headline was blunt: "Building on top of a Chinese model feels particularly fraught right now"
Moonshot AI cheerfully congratulated Cursor on Twitter, calling it "the open model ecosystem we love to support"
Community reaction was split:
- One camp sees this as open source working as intended
- The other worries about geopolitical risk as US-China AI decoupling accelerates
What this means for us: A reminder about supply chain risk in AI tooling. Cursor is many developers' primary IDE, but model provenance affects enterprise trust. For security-sensitive projects, tool choices need clear model source documentation
Flash-MoE: 397B Model on a Laptop
Apple's "LLM in a Flash" research was validated by Dan Woods — he got Qwen3.5-397B-A17B running at 5.5+ tokens/second on a 48GB MacBook Pro M3 Max
The key technique is MoE (Mixture of Experts): 397B total parameters, but only 17B active per inference. With quantization, it needs 120GB (using NVMe swap to bridge the memory gap)
The same week, Hacker News featured tinybox (a compact deep learning computer) and professional video editing running entirely in-browser with WebGPU + WASM
What this means for us: Local LLM viability keeps improving. For privacy-sensitive projects (healthcare, legal), this is an important deployment option. Both Med Vision and TFT may need "data stays on-premise" architectures
OpenAI Building a Fully Automated Researcher
MIT Technology Review reports OpenAI is refocusing its research direction — the goal is to build a fully automated AI researcher
This isn't a regular chatbot. It's an agent system capable of autonomously tackling large, complex problems. Simultaneously, OpenAI is discussing plans with the Pentagon to let AI companies train military-specific models on classified data
The same week, OpenAI published a paper on monitoring internal coding agents using chain-of-thought analysis to detect misalignment
What this means for us: OpenAI's direction is shifting from "conversational tool" to "autonomous agent." This aligns with how we use Claude Code — agents don't just answer questions, they autonomously complete work. The misalignment monitoring methods are worth studying
Dreamer: Personal Agent OS
Latent Space reported that /dev/agents has officially rebranded as Dreamer, built by former Android VP David Singleton
The vision is a "Personal Agent Operating System" — not just a single agent, but an entire agent ecosystem runtime
Latent Space is offering $10,000 prizes for new tools, with special access for subscribers
What this means for freelancers: The Personal Agent OS concept parallels the ops system I've been building — a central system managing multiple project agents. The difference: Dreamer is a general-purpose platform; mine is workflow-specific for freelance project management. Worth tracking
Simon Willison: Using Git with Coding Agents
Simon Willison published an important practical guide: how to use Git effectively with coding agents
Core insight: Git isn't just version control — it's the safety net for agents. Because agents make mistakes, Git lets you track changes and reverse errors
The same week, he wrote about Starlette 1.0 (the framework underlying FastAPI), calling it severely underrated in the Python ecosystem
Another interesting experiment: using AI to profile Hacker News users based on their comments. He called it "mildly dystopian"
What this means for us: Our Git workflow (worktree isolation, pre-commit hooks, branch strategy) is already aligned with this direction. The Starlette 1.0 update is also notable since our backends all run FastAPI
LangChain Fleet + Sandboxes
LangChain shipped multiple releases this week:
| Product | Function |
|---|---|
| Fleet (formerly Agent Builder) | Enterprise agent management platform |
| Sandboxes | Secure code execution environments for agents |
| Open SWE | Open-source internal coding agent framework |
| Deploy CLI | Command-line agent deployment to LangSmith |
LangChain's direction is clear: shifting from "framework" to "platform," competing for enterprise agent infrastructure
What this means for us: Open SWE uses LangGraph + Deep Agents architecture, aligned with the agent orchestration patterns we've researched. Sandboxes are a practical reference for security-sensitive projects
Iran War Threatens AI Chip Supply Chain
Financial Times' most important piece this week: "How the Iran war could derail the AI boom"
Core argument: the entire chip supply chain depends on energy and chemical imports from the Middle East. The war is now in its fourth week, with Strait of Hormuz closure threats escalating
| Indicator | Status |
|---|---|
| Oil prices | Most volatile in 40 years |
| Airlines | Already preparing for oil crisis |
| Gold | Rebounding after worst weekly drop in 40 years |
| LNG supply | Last Middle East shipments arriving within 10 days |
Bloomberg simultaneously reports Australian stocks nearing technical correction, with Singapore bonds emerging as a safe haven
What this means for us: If the war continues to escalate, AI hardware costs could rise. Short-term impact on our freelance work is minimal, but long-term cloud service pricing (GCP, AWS) bears watching. Taiwan's semiconductor supply chain may also face indirect effects
Action Items
- Track GitHub Squad — Our multi-agent system can learn from its inspectable/predictable design principles
- Study Open SWE — LangGraph + Deep Agents architecture as reference for TEMC RAG pipeline
- Monitor toolchain supply chain — Cursor/Kimi incident reminds us: enterprise clients care about model provenance
- Local LLM progress — Flash-MoE makes 400B models laptop-viable; relevant for healthcare/legal deployments
- Watch energy geopolitics — Iran war's impact on AI infrastructure is materializing
Sources
- How Squad runs coordinated AI agents inside your repository — GitHub Blog
- Cursor admits its new coding model was built on top of Moonshot AI's Kimi — TechCrunch
- Autoresearching Apple's "LLM in a Flash" — Simon Willison
- OpenAI is throwing everything into building a fully automated researcher — MIT Technology Review
- Dreamer: the Personal Agent OS — Latent Space
- Using Git with coding agents — Simon Willison
- Introducing LangSmith Fleet — LangChain Blog
- How the Iran war could derail the AI boom — Financial Times
- How we monitor internal coding agents for misalignment — OpenAI
- Experimenting with Starlette 1.0 — Simon Willison
- Open SWE: An Open-Source Framework for Internal Coding Agents — LangChain Blog
