Sift.

Week 2026-29 · Jul 13–19, 2026

10 stories · 20/20 feeds live · $0.28 run

Week at a glance

Models & Research 6Tooling 3Policy 1

hi 9 · lo 5 · avg 6.8

Models & Research

9
#1

Moonshot AI releases Kimi K3, a 2.8T-parameter open-weight model claiming frontier coding performance

Moonshot AI announced Kimi K3, described as a 2.8-trillion-parameter (A50B) model and the largest open model to date, with an open-weight release promised by July 27, 2026. Self-reported benchmarks show it competitive with or beating Claude Opus/Fable 5 and GPT-5.6 on agentic coding at roughly one-third the price.

Major open model release with concrete specs and benchmarks, directly relevant to model capabilities and open-weight ecosystem.

8
#2

Thinking Machines releases Inkling, a 975B open-weights multimodal model

Thinking Machines Lab released Inkling, an Apache-2.0 licensed Mixture-of-Experts multimodal transformer with 975B total (41B active) parameters trained on 45 trillion tokens. A smaller 276B (12B active) variant, Inkling-Small, was announced as forthcoming pending further testing.

First major open-weights release from Mira Murati's lab with detailed specs; strong interest for model releases and open weights.

7
#5

NVIDIA and Stanford introduce RoboTTT, scaling robot policies to 8K-timestep context via test-time training

NVIDIA GEAR Lab and Stanford introduced RoboTTT, scaling robot visuomotor context to 8,000 timesteps at constant inference cost using test-time training that updates a small internal model per sensor reading. They report a context-scaling curve where 8K-context pretraining beats 1K by 62% with one-shot imitation from human video and mid-episode error recovery.

Capabilities research with concrete context-scaling numbers and a primary paper/blog; relevant to research interest.

6
#7

GPT-5.6 used in convex optimization and hard-problem experimentsneeds verification

A widely-discussed Reddit thread claims GPT-5.6 helped close a 30-year gap in convex optimization via a prompt, while other posts benchmark GPT-5.6 Sol against Fable 5 on NP-hard problems. Separate demos show GPT-5.6 building a SQLite-based Doom-like game engine.

Frontier model capability claims with concrete tasks, though sourced from Reddit/blogs rather than primary papers.

6
#9

OpenAI details GPT-Red red-teaming; Claude web_fetch exfiltration hole found

OpenAI described GPT-Red, an automated red-teaming system using self-play to improve robustness against prompt injection. Separately, a researcher demonstrated a data-exfiltration vulnerability in Claude's web_fetch tool via the lethal-trifecta pattern.

Concrete safety/robustness research and a real agent security vulnerability; relevant to evals and agent safety.

5
#10

Research and engineering notes on reasoning effort, model routing, and agent building

Sebastian Raschka analyzes how LLMs learn low-, medium-, and high-effort reasoning modes, while IBM Research examines the complexities of model routing. Additional posts cover fine-tuning diffusion models at scale with NVIDIA NeMo and lessons from building the Shippy agent.

Useful technical research and engineering writeups on reasoning control and agent infrastructure.

Tooling

8
#3

Codex usage reportedly up 10x in 6 months to 7M usersneeds verification

AINews reports OpenAI's Codex grew more than 10x in six months to roughly 7M users, adding over 1M in a single day, raising questions about whether it overtook Claude Code. A related 'Codex Resets' discussion trended on Hacker News amid the usage surge.

Concrete adoption numbers on coding agents, central to the reader's agentic SDLC interest.

7
#4

Grok CLI caught uploading local files to the cloud; xAI open-sources grok-build

xAI's Grok CLI faced backlash after users found it could upload entire working directories—including SSH keys and password databases—to xAI's cloud buckets, after which xAI open-sourced the tool as grok-build. Separately, a Codex bug was documented in which the agent could delete a user's $HOME directory under full-access mode.

Significant dev-tooling security incident with concrete details on an agent harness; relevant to coding agents and security.

6
#8

Context and 'loop' engineering emerge as core practices for agentic development

Practitioners including Dex Horthy and Armin Ronacher discuss context engineering, 'loop engineering,' and how coding agents fit into human organizations and shared project understanding. Separately, Claude Code was confirmed to now ship on a Rust port of Bun, and GitHub's Dependabot added default dependency cooldowns.

Directly relevant to agentic SDLC and dev tooling practices, with primary practitioner commentary.

Policy

6
#6

Debate intensifies over open-weight AI viability and US policyneeds verification

Following the Kimi K3 and Inkling releases, commentators including Yann LeCun and industry figures argued open-weight models are essential for economic competitiveness, citing token-cost disparities of $26-56 versus $0.50-1 per million tokens. Interconnects framed the moment as a critical test of open-source AI's viability over the next six months.

Substantive policy discussion on open models with cost figures, tied to open-weight ecosystem the reader tracks; some punditry lowers it.

Sources scanned · 20/20 live