THE AI PULSEEN

The Pulse — March 20, 2026

The signals that entered our radar, organized with sources and context to understand what changed.

ModelsAgentsAnthropic
LISTEN TO THIS EDITION

The audio script is ready; narration will appear after voice generation finishes.

  1. 01Hugging Face (NVIDIA) + arXiv

    Nemotron-Cascade 2 (30B MoE, 3B active) — model + paper drop

    WHY IT ENTERED THE RADAR

    This is the “intelligence density” story: frontier-ish math/code/agentic behavior with 3B activated params. Also a concrete recipe: Cascade RL + multi-domain on-policy distillation.

    SUGGESTED EDITORIAL ANGLE

    “The new open model that ‘acts big’ while staying small: what MoE + post-training actually bought them (and what it didn’t).”

    Open original source ↗
  2. 02arXiv

    Doc-to-LoRA (D2L): “internalize a long doc instantly” by generating a LoRA adapter

    WHY IT ENTERED THE RADAR

    A clean way to reframe long-context costs: instead of paying quadratic attention repeatedly, compile a document into a tiny adapter in one forward pass, then query cheaply.

    SUGGESTED EDITORIAL ANGLE

    “Long-context is expensive—what if you compile documents into LoRAs on the fly?” (Use a simple mental model: ‘KV cache vs. adapter cache’.)

    Open original source ↗
  3. 03Anthropic newsroom

    Anthropic: “Detecting and preventing distillation attacks” (DeepSeek / Moonshot / MiniMax)

    WHY IT ENTERED THE RADAR

    Upstream, concrete numbers (claims of 16M exchanges / 24k accounts) + a real playbook of how labs attempt capability extraction. Also touches export-controls narrative: ‘apparent rapid progress’ vs ‘borrowed outputs.’

    SUGGESTED EDITORIAL ANGLE

    “Distillation is normal… until it isn’t: the line between ‘compression’ and ‘capability theft’ + what this means for open weights.”

    Open original source ↗
  4. 04Claude Code docs

    Claude Code “Channels”: push events into a running agent session (Telegram/Discord in preview)

    WHY IT ENTERED THE RADAR

    This is an architecture shift: agents aren’t only pull-based (cron/scheduled). They become reactive systems (inbound events) while the session is alive—closer to ‘always-on ops’.

    SUGGESTED EDITORIAL ANGLE

    “Agents that wake up when something happens: how event-driven ‘channels’ change automation vs. scheduled polling.”

    Open original source ↗
  5. 05Claude product blog

    Claude now creates interactive charts/diagrams/visualizations inline (beta)

    WHY IT ENTERED THE RADAR

    Visuals as ephemeral conversational UI (not ‘Artifacts’). This suggests LLM UX is becoming a live notebook-like surface, not just text + attachments.

    SUGGESTED EDITORIAL ANGLE

    “The new battleground is UI: chat is turning into mini-apps—what creators should do with interactive visuals (and how it’ll be faked by ‘AI tool’ aggregators).”

    Open original source ↗
  6. 06Google DeepMind model page

    Gemini 3.1 Flash-Lite: “scalable thinking model” for high-volume, low-latency work

    WHY IT ENTERED THE RADAR

    This is the ‘production model’ story: throughput + structured output compliance + selectable thinking level. The table also reveals where the market is heading: price/speed as first-class metrics.

    SUGGESTED EDITORIAL ANGLE

    “Why 2026 is the year of fast reasoning: the economics of ‘good enough’ models that win on latency + cost.”

    Open original source ↗
  7. 07SkyPilot blog

    Scaling Karpathy’s “autoresearch” with a GPU cluster (parallel agent search)

    WHY IT ENTERED THE RADAR

    Great upstream case study: parallelism changes agent behavior (factorial experiment waves vs greedy hill-climbing). This is one of the clearest “agent + infra” pieces you can adapt into creator content.

    SUGGESTED EDITORIAL ANGLE

    “Agents don’t just get faster with more GPUs—they get smarter search. Here’s why parallelism changes the optimization strategy.”

    Open original source ↗
  8. 08GitHub

    KittenTTS (v0.8): high-quality TTS on CPU with tiny models (as low as ~25MB int8)

    WHY IT ENTERED THE RADAR

    Edge TTS is sneaking up: small ONNX models that run without GPU unlock offline voice features in apps, agents, and devices—without paying API costs.

    SUGGESTED EDITORIAL ANGLE

    “The underrated trend: tiny voice models on CPU. Where this beats cloud TTS (latency, privacy, cost).”

    Open original source ↗
  9. 09Science (via HN)

    arXiv declares independence from Cornell (platform governance / funding)

    WHY IT ENTERED THE RADAR

    This affects the most upstream distribution channel in ML research. Governance changes can influence moderation, sustainability, and product direction (APIs, metadata, partnerships).

    SUGGESTED EDITORIAL ANGLE

    “Why arXiv governance matters for AI: the supply chain of research discovery (and how creators can monitor new papers faster).”

    Open original source ↗
  10. 10Original source

    Creator-watch (new uploads worth scanning, then go upstream)

    Open original source ↗
TAKE THIS PULSE TO YOUR AI

Continue the analysis where you already work.

Copy this prompt into ChatGPT, Claude, Gemini, or whichever AI you use. It includes the signals, sources, and a guide for turning them into decisions.

No account is connected and no data is shared automatically.
PROMPT.md