LLM
15 post tagged with this.
TypeSafe AI has launched Jev, the first public System One foundation model. Instead of generating text, it outputs typed decisions, calibrated probabilities, and confidence scores in 70ms. A deep dive into RLCD and agent architectures.
A comprehensive technical breakdown of the Google Gemini model family. Explore the differences between Gemini Flash and Pro, 2M+ context windows, Context Caching, and the Multimodal Live API.
The AI industry has shifted from brute-force pre-training scaling laws to test-time compute. A deep dive into inference scaling, Process Reward Models (PRMs), System 2 reasoning, and token economics.
xAI announced Grok 4.6, optimized for coding, agentic tasks, and knowledge work. A 500,000-token context window, and the date Elon Musk gave for Grok 4.7.
Anthropic announced Claude Opus 5: it lands within 0.5 points of Fable 5 on CursorBench while costing half as much. Here's the five-tier effort system, the pricing, and its new role in Claude Max.
Meta Superintelligence Labs shipped Muse Spark 1.1 to catch up with Anthropic and OpenAI in coding. A 1-million-token context window, multi-agent orchestration, and benchmarks that beat Gemini.
OpenAI's new model family GPT-5.6 became the first AI model to clear a White House review before shipping. Here's how Sol, Terra, and Luna differ, and how the rollout unfolded.
xAI shipped Grok 4.5, a 1.5-trillion-parameter model trained on real-world coding data from Cursor. Here's the pricing, the capabilities, and how it stacks up against the competition.
Anthropic just shattered AI benchmarks with the release of Claude Fable 5 and Mythos 5. From massive codebase migrations to autonomous scientific discovery, the 'Mythos-class' is redefining what AI can do.
From managing the context window and writing effective CLAUDE.md files to running parallel sessions and automation pipelines — proven patterns for maximizing Claude Code productivity.
DeepSeek released V4-Pro and V4-Flash on April 24, 2026 — both open-source, MIT licensed, with 1M token context windows. V4-Pro tops LiveCodeBench at 93.5% and costs 7x less than Claude Opus. V4-Flash undercuts GPT-5.4 Nano on price. Legacy models retire July 24.
OpenAI launched GPT-5.5 on April 23, 2026 — six weeks after GPT-5.4. The model delivers better results with fewer tokens, a 400K context window in Codex, and meaningful gains on scientific research and agentic coding. API pricing is 2x GPT-5.4. API access is coming soon.
Anthropic shipped Claude Opus 4.7 today with real benchmark data against GPT-5.4 and Gemini 3.1 Pro. 87.6% on SWE-bench Verified, 80.6% on document reasoning, 3.75MP vision, and a new xhigh effort level.
KV Cache explained: Discover how Key-Value (KV) Cache works in large language models (LLMs), how self-attention uses it, why it drops inference complexity from O(n²) to O(n), and how prompt caching slashes API costs by up to 90%.
TurboQuant compresses LLM key-value caches to 3 bits with no accuracy loss — 8x throughput on H100 GPUs and zero training required.