LLM

12 post tagged with this.

xAI announced Grok 4.6, optimized for coding, agentic tasks, and knowledge work. A 500,000-token context window, and the date Elon Musk gave for Grok 4.7.

Anthropic announced Claude Opus 5: it lands within 0.5 points of Fable 5 on CursorBench while costing half as much. Here's the five-tier effort system, the pricing, and its new role in Claude Max.

Meta Superintelligence Labs shipped Muse Spark 1.1 to catch up with Anthropic and OpenAI in coding. A 1-million-token context window, multi-agent orchestration, and benchmarks that beat Gemini.

OpenAI's new model family GPT-5.6 became the first AI model to clear a White House review before shipping. Here's how Sol, Terra, and Luna differ, and how the rollout unfolded.

xAI shipped Grok 4.5, a 1.5-trillion-parameter model trained on real-world coding data from Cursor. Here's the pricing, the capabilities, and how it stacks up against the competition.

Anthropic just shattered AI benchmarks with the release of Claude Fable 5 and Mythos 5. From massive codebase migrations to autonomous scientific discovery, the 'Mythos-class' is redefining what AI can do.

From managing the context window and writing effective CLAUDE.md files to running parallel sessions and automation pipelines — proven patterns for maximizing Claude Code productivity.

DeepSeek released V4-Pro and V4-Flash on April 24, 2026 — both open-source, MIT licensed, with 1M token context windows. V4-Pro tops LiveCodeBench at 93.5% and costs 7x less than Claude Opus. V4-Flash undercuts GPT-5.4 Nano on price. Legacy models retire July 24.

OpenAI launched GPT-5.5 on April 23, 2026 — six weeks after GPT-5.4. The model delivers better results with fewer tokens, a 400K context window in Codex, and meaningful gains on scientific research and agentic coding. API pricing is 2x GPT-5.4. API access is coming soon.

Anthropic shipped Claude Opus 4.7 today with real benchmark data against GPT-5.4 and Gemini 3.1 Pro. 87.6% on SWE-bench Verified, 80.6% on document reasoning, 3.75MP vision, and a new xhigh effort level.

A deep dive into Key-Value (KV) Cache in large language models — what it is, how attention uses it, when it activates, and how it reduces latency and API costs.

TurboQuant compresses LLM key-value caches to 3 bits with no accuracy loss — 8x throughput on H100 GPUs and zero training required.

← All posts