Sift.

Week 2026-33 · Aug 10–16, 2026

10 stories · 20/20 feeds live · $0.25 run

Week at a glance

Models & Research 6Tooling 2Infra 1Business 1

hi 8 · lo 4 · avg 5.9

Models & Research

8
#1

Google DeepMind releases Gemini 3.7 Flash

Google DeepMind introduced Gemini 3.7 Flash, adding new capabilities to its fast model tier alongside earlier Flash variants. Same-day plugin support (llm-gemini 0.33) exposed reasoning traces and server-side tools, and analysts framed the release as returning GDM to frontier competitiveness.

Major model release from a frontier lab with same-day tooling support.

8
#2

Meta returns to open weights with Muse Glimmer (30B, Apache 2.0)

Meta released Muse Glimmer, a 30B dense multimodal model under Apache 2.0 that runs on a single GPU and targets agentic task completion. Independent benchmarks put it at 35 on the Artificial Analysis Intelligence Index with strong tool use but a high hallucination rate, marking Meta's first open release since Llama 4.

Significant open-weights release with concrete benchmarks and single-GPU deployment relevance.

7
#3

Paper: stealing reasoning traces from proprietary LLM APIs

A paper shows that encrypted chain-of-thought blocks returned by Anthropic, OpenAI, and Google can be replayed across sessions and into weaker sibling models to recover hidden reasoning in plaintext. The technique combines replay with jailbreaking of the weaker model.

Novel security research affecting all major API providers and agent security.

6
#5

Chinese and open models track the frontier: DeepSeek V4 Pro, GLM-5.3needs verification

DeepSeek released V4 Pro 0813 via API with weights later posted to Hugging Face at 1.7T parameters, while analysis argued GLM-5.3 shows Chinese labs keeping pace with the frontier beyond distillation. Hugging Face's summer 2026 report surveys the broader open-model landscape.

Frontier open-weight releases with parameter counts, relevant to model landscape.

5
#7

Grok 4.6 and Grok Bot enter the AI-teammate/agent categoryneeds verification

Latent Space and The Pragmatic Engineer covered SpaceXAI's Grok 4.6 and Grok Bot as a significant new entrant in the managed AI-agent/teammate category. Commentary framed Grok Bot as a potential inflection moment for autonomous agents.

New agent-oriented entrant, borderline product but relevant to agentic tooling.

5
#8

Reproducing 2,200 ICML papers: lessons in eval reproducibility

Hugging Face reported on an effort to reproduce 2,200 papers from ICML and shared findings about what did and did not replicate. The work documents systemic reproducibility challenges in ML research.

Evals/reproducibility data relevant to research rigor interest.

Tooling

7
#4

Agentic coding harnesses mature: Claude Code auto mode default, multi-agent patterns

Anthropic is making auto mode the default in Claude Code across Pro, Max, and Team plans and published research on patterns and problems in emerging multi-agent systems. Separately, Latent Space profiled Flue 2, a React-inspired meta-harness for defining agents through hooks.

Directly on agent harnesses and SDLC automation, the reader's stated focus.

5
#6

GitHub Models service retired

GitHub has completed the retirement of GitHub Models, its unified multi-provider LLM API and playground. Existing GitHub Actions workflows relying on the service now fail.

Dev-tooling infrastructure change affecting CI/LLM workflows.

Infra

4
#9

OpenAI previews Ultrafast mode: GPT-5.6 Sol at up to 14x speed via Cerebras

OpenAI previewed an Ultrafast API tier that runs GPT-5.6 Sol up to 14x faster, delivering up to 750 output tokens per second. The service is powered by Cerebras hardware.

Concrete inference-infra numbers (750 tok/s, Cerebras) matching compute interest.

Business

4
#10

OpenAI enterprise adoption research and agentic deployments (Codex, ChatGPT Work)

OpenAI published research on how enterprises adopt agentic AI via ChatGPT and Codex, alongside case studies from RingCentral, Virgin Atlantic, Zapier, and Model ML. The materials describe concrete workflow automation across engineering, finance, marketing, and operations.

Real enterprise deployments and adoption data align directly with reader's core interest.

Sources scanned · 20/20 live