August 20, 2026 · Thursday

The Model Rush08.20
Infra · NVIDIA

CoreWeave: Vera Rubin Lifts Per-Megawatt Throughput 10×

CoreWeave measured a 10× jump in tokens per second per megawatt on NVIDIA's Vera Rubin platform, underscoring its energy-efficiency edge for inference.

Models · Zhipu

GLM-5.3 Ships, Beating DeepSeek V4-Pro on Terminal-Bench

Cline's evaluation has GLM-5.3 ahead of Fable and yesterday's DeepSeek V4-Pro 0813 on Terminal-Bench, with standout coding performance.

OSS · Hugging Face

Sentence Transformers v6 Centers on Multi-Vector Embeddings

The new major release pivots to multi-vector embedding models, a shift that could lift retrieval and representation quality.

Models · Alibaba

Qwen3.8-27B Reaches #1 on Cline

Community builders pushed Qwen3.8-27B to first among open-weight models on Cline in four days.

Video · MiniMax

A Streaming LoRA Adapter for MiniMax H3 Appears

The community built a streaming video LoRA adapter for H3. MiniMax says real-time generation still needs inference acceleration — and floats an online hackathon.

Models · xAI

Grok 4.6 Lands on Amazon Bedrock

xAI's latest model is now available through Amazon Bedrock, expanding its enterprise distribution.

OSS · Model Release

Ornith-1.5 Ships a 9B/35B/397B Open LLM Family

A family of open-source LLMs spans 9B dense, 35B MoE and 397B MoE, trained with self-improvement.

Insight · GLM

GLM-5.3: No New Base, Pure Post-Training

Zhipu kept the GLM-5.2 base — a 743B-parameter MoE activating roughly 40B — and spent a month on long-horizon RL to lift coding by 50%.

OSS · Apache

Apache Incubates Its First Agent Harness, Maka

Maka, a local-first AI desktop assistant, enters the Apache Incubator as the foundation's first agent-harness project.

Also NotedSignals

© 2026 FAV0 · AI Daily · The day's AI news, set in type