Key Shifts

  • OpenAI unveils first custom inference chip “Jalapeño” with Broadcom: OpenAI revealed its first custom-built inference processor, designed to reduce dependence on NVIDIA GPUs. This puts OpenAI on the same path as Google (TPU) and Amazon (Trainium). The chip is optimized specifically for LLM inference, and OpenAI’s own models assisted in its design. Broadcom handled manufacturing. TechCrunch · OpenAI
  • Qualcomm acquires Modular for ~$4B: Qualcomm is buying Modular, the AI infrastructure startup founded by Chris Lattner (creator of Swift and LLVM). Modular built Mojo — a Python superset for high-performance AI — and the MAX platform. This is a strategic bet to own the AI software stack for on-device and edge inference. AI infrastructure M&A is accelerating. Modular · Qualcomm
  • Google bakes Computer Use into Gemini 3.5 Flash: Google integrated computer control as a built-in tool in Gemini 3.5 Flash, enabling clicking, typing, and UI navigation across browser, mobile, and desktop environments. Previously this was only available as a standalone Gemini 2.5 model. Available via Gemini API and Enterprise Agent Platform. Google Blog

Startup / Product / Platform Radar

  • Reid Hoffman calls xAI a “complete train wreck”: In a candid Fortune interview, the LinkedIn co-founder and influential AI investor assessed the AI landscape. He argued SpaceX is not an AI company, called xAI a “complete train wreck,” and said there is still room for OpenAI and Anthropic. Worth noting as an insider’s frank map of the competitive landscape. Fortune

AI Future Signals

  • Oracle cuts 21,000 jobs while pouring billions into AI infrastructure: Oracle is simultaneously laying off 21,000 employees and investing heavily in AI CapEx. The pattern of headcount reduction funding AI infrastructure is becoming a template for large enterprises. This matters for founders building tools that replace or augment enterprise workflows. Ars Technica
  • Meta pauses AI employee-tracking program after internal data leak: Meta had deployed AI-based employee activity monitoring. An internal breach exposed the data, and the program is now paused. A concrete case of AI-powered corporate surveillance hitting real organizational friction — the tension between productivity measurement and privacy is escalating. WIRED
  • Qwen-AgentWorld: Language World Models for General Agents: Alibaba’s Qwen team published a paper on arXiv introducing language models that act as world simulators for AI agents. Trained on 10M+ real-world interaction trajectories across 7 domains, the models come in two sizes (35B-A3B and 397B-A17B) and use a three-stage training pipeline: CPT, SFT, and RL. LLMs simulating environment dynamics instead of explicit world models is a notable architectural direction. arXiv

Realistic Opportunities / Experiments

  • Revisit cost-prohibitive AI products under the inference cost curve: As custom inference silicon from OpenAI (and others) reaches production scale, LLM inference costs will likely drop further. Products currently uneconomical due to inference expense — real-time large-scale document processing, high-frequency agent loops — are worth recalculating on a 6–12 month horizon.
  • Rethink agent UX now that Computer Use is table stakes: When even mid-tier models ship with built-in computer control, the real question for founders becomes: is handing a GUI to an agent actually the UX users want, or is tight API/shortcut integration better? The answer will be domain-specific — start experimenting now.

Uncertainties / Keep Watching

  • How fast will custom AI chips erode NVIDIA’s inference dominance?: OpenAI, Google, Amazon, and Microsoft are all accelerating custom silicon strategies. But NVIDIA continues to ship new inference-optimized generations annually. The pace and magnitude of market share shift remain unclear until production volumes and independent benchmarks arrive. TechCrunch