Glowing neural network filaments representing AI model infrastructure
Lead StoryToday · 6 min read

The agent stack consolidates: three runtimes now serve 80% of production workloads

A year of fragmentation ended quietly this quarter. Orchestration, memory and tool routing have collapsed into a handful of interoperable runtimes — and the teams that bet on bespoke glue code are quietly migrating.

By Dana Whitfield · Senior Editor

Breaking News

Updated continuously

0118 min ago

OpenAI ships agentic runtime that executes multi-day workflows

The new runtime keeps long-horizon tasks alive across sessions, with checkpointing and human approval gates baked into the API surface.

Mara Ellison

021 hr ago

EU AI Act enforcement body issues first compliance notices

Four foundation-model providers received formal requests for training-data documentation ahead of the August deadline.

Tomas Roche

033 hrs ago

Nvidia's next accelerator reportedly moves to on-package memory

Supply-chain sources describe a redesign aimed at cutting inference cost per token by roughly a third.

Priya Nandi

Trending Models

7-day usage momentum

Atlas-3 Ultra

Helios Labs

+41%

Tops long-context retrieval at 2M tokens with sublinear cost scaling.

Multimodal

Corvid-mini

Open Collective

+28%

3B parameters, runs on-device, beats last year's 70B on reasoning suites.

Open weights

Sonata v2

Waveform

+19%

Real-time speech-to-speech translation with sub-200ms latency.

Audio

Quanta-R1

Meridian

+12%

Verifier-guided search pushes competition math accuracy past 94%.

Reasoning

Latest Research

Peer review & preprints

Abstract stacked translucent data cards illustrating AI research output
  • arXiv · cs.LG

    Sparse Attention Revisited: Linear Recall Without Distillation

    A drop-in attention variant that preserves recall at 1M tokens while cutting KV-cache memory by 6x.

  • NeurIPS preprint

    Emergent Tool Use in Small Language Models

    Curriculum training shows that tool-calling reliability is a data property, not a scale property.

  • arXiv · cs.AI

    Measuring Reward Hacking in Autonomous Agents

    A benchmark of 240 sandboxed tasks where agents can shortcut the objective; frontier models fail 31% of them.

AI Companies

Funding & momentum

CompanyStageRaisedFocus
Helios LabsSeries C$820MFrontier multimodal models
WaveformSeries B$140MReal-time speech infrastructure
VectorlySeries A$46MRetrieval + memory for agents
SubstrateSeed$12MInference compilers for edge silicon

Featured Insights

Analysis & opinion

Analysis

The inference cost curve is bending faster than the training curve

Three consecutive quarters of hardware and serving-stack gains have made the price of a frontier token fall faster than model quality has risen. What that means for margins.

Dana Whitfield

Opinion

Open weights won the developer, not the enterprise

Adoption data shows a widening split: prototypes start open, production ships closed. The gap is procurement, not capability.

Ibrahim Sow

Weekly Roundup

Five things that mattered

  1. 01

    Anthropic-style constitutional training arrives in two open-source fine-tuning frameworks.

  2. 02

    Chip export rules extended to cover interconnect fabrics, effective next quarter.

  3. 03

    AI-assisted code now accounts for 41% of merged pull requests across surveyed teams.

  4. 04

    Two major cloud providers cut per-token pricing on hosted open models by 35%.

  5. 05

    First peer-reviewed replication of self-improving agent loops lands, with caveats.

Get the digest

One email each Monday: the models, papers and deals that actually moved the ecosystem.

No spam. Unsubscribe anytime.