THE AI PULSEEN

The Pulse — May 17, 2026

The signals that entered our radar, organized with sources and context to understand what changed.

ModelsAgentsOpenAI
LISTEN TO THIS EDITION

The audio script is ready; narration will appear after voice generation finishes.

  1. 01Thinking Machines Labs (primary blog)

    Interaction Models (real-time, micro-turn, multi-stream human–AI collaboration)

    WHY IT ENTERED THE RADAR

    This is a concrete architectural proposal for “real-time collaboration” models (simultaneous perception + output) instead of bolting interactivity on with harnesses (VAD, turn-taking, etc.). It also explicitly frames interactivity as a scaling axis alongside intelligence.

    SUGGESTED EDITORIAL ANGLE

    “The next LLM jump isn’t bigger context — it’s time (micro-turns, interruption, overlap). Here’s what changes when models are trained for interaction natively.”

    Open original source ↗
  2. 02OpenAI (primary)

    Codex “from anywhere” (mobile steering + remote connections + hooks)

    WHY IT ENTERED THE RADAR

    The product direction is “agents that run long” + human check-ins as the core UX. The important detail is the workflow primitives: remote connections, approvals, and enterprise hooks (validators, secret scanning, logging, memories).

    SUGGESTED EDITORIAL ANGLE

    “The real story isn’t ‘Codex on your phone’ — it’s the emerging agent operating model: approvals, hooks, and remote sessions as the new CI/CD.”

    Open original source ↗
  3. 03arXiv (primary paper)

    $δ$-mem: efficient online memory for LLMs (delta-rule state + low-rank attention corrections)

    WHY IT ENTERED THE RADAR

    A clean middle path between ‘just increase context’ and ‘full fine-tuning’: keep the backbone frozen, add a tiny online associative memory state (e.g., 8×8) that writes continuously and modulates attention. This is exactly the kind of mechanism agent builders can watch for in next-gen “memory-native” models.

    SUGGESTED EDITORIAL ANGLE

    “You can give an LLM memory with an 8×8 matrix — here’s the trick, and what it implies for assistants that learn over time.”

    Open original source ↗
  4. 04r/LocalLLaMA (community report; practical upstream signal)

    llama.cpp: MTP (multi-token prediction) support testing on Qwen3.6 + RTX 5090

    WHY IT ENTERED THE RADAR

    Local inference speed is still the biggest gating factor for “always-on” assistants. MTP can change perceived latency dramatically (especially for short outputs), and community benchmarks show where it helps and what flags matter.

    SUGGESTED EDITORIAL ANGLE

    “The fastest path to a better ‘assistant feel’ might be MTP + good decoding, not new models. Here’s what to toggle in llama.cpp and what to expect.”

    Open original source ↗
  5. 05NVIDIA Labs (primary project page)

    NVIDIA SANA-WM (minute-scale world modeling; 720p video claim)

    WHY IT ENTERED THE RADAR

    World models + longer horizon generation is creeping from research into “almost productizable” claims (minute-scale, 720p). Even without full details in the page fetch, it’s a strong upstream breadcrumb worth tracking for releases (paper/code) and the next wave of video generation tooling.

    SUGGESTED EDITORIAL ANGLE

    “World models are quietly eating video generation: what ‘minute-scale world modeling’ could enable (consistent scenes, controllable camera, long actions).”

    Open original source ↗
  6. 06OpenAI (primary)

    New personal finance experience in ChatGPT (account linking + financial memories)

    WHY IT ENTERED THE RADAR

    This is a template for how “agentic” features will ship to normal users: connect to real systems (Plaid), add domain-specific memories, add stronger guardrails + deletion controls. It’s also a distribution wedge for deeply context-grounded assistants.

    SUGGESTED EDITORIAL ANGLE

    “Finance is the first mainstream ‘connected agent’ category. What it teaches us about data permissions, memory, and trust UX.”

    Open original source ↗
  7. 07Anthropic newsroom (primary)

    Anthropic Labs: Claude Design (visual work: slides, prototypes, one-pagers)

    WHY IT ENTERED THE RADAR

    Labs products are where interaction patterns show up first. “Design” is a good proxy for where multimodal + tool use is heading: structured outputs, iteration loops, and publishing-ready artifacts.

    SUGGESTED EDITORIAL ANGLE

    “AI is moving from ‘generate’ to ‘ship’. Design workflows are the canary — because the output quality is instantly obvious.”

    Open original source ↗
  8. 08HN link to crates.io (upstream)

    Zerostack: Unix-inspired coding agent written in Rust

    WHY IT ENTERED THE RADAR

    The agent ecosystem is fragmenting: more local-first, compiled, permissioned agents instead of one monolithic cloud assistant. Rust-based agents are interesting for sandboxing and distribution.

    SUGGESTED EDITORIAL ANGLE

    “Why coding agents are turning into Unix tools: composable, local, strict permissions — and what that means for dev workflows.”

    Open original source ↗
TAKE THIS PULSE TO YOUR AI

Continue the analysis where you already work.

Copy this prompt into ChatGPT, Claude, Gemini, or whichever AI you use. It includes the signals, sources, and a guide for turning them into decisions.

No account is connected and no data is shared automatically.
PROMPT.md