Loading...
15 entries with this tag
On August 14, 2026, Alibaba's Qwen team released Qwen3.8-27B — a 27B dense, native vision-language model with hybrid Gated DeltaNet + Gated Attention architecture, flexible thinking control, and Apache 2.0 licensing. The model delivers 73.0 on Terminal Bench 2.1 (within 5 points of Opus 4.6 Max), 61.7 on SWE-bench Pro, 84.3 on OSWorld-Verified, and 90.0 on MathVision, all in a model that fits on a single consumer GPU. Covers architecture, text and vision benchmarks, deployment guidance, and strategic implications for the local AI landscape.
Alibaba released Qwen3.8-Max on August 3, 2026 — a 2.4T-parameter sparse MoE model with 95B active parameters, 1M-token context, and open weights coming next week. Covers the architecture, benchmark results (86.6 Terminal-Bench 2.1, 56.6 DeepSWE 1.1, 73.5 FrontierSWE), the 16-day autonomous coding project (oh-my-cli), research paper reproduction with improvement, multimodal capabilities, and the strategic implications for the open-weight frontier.
Moonshot AI releases Kimi K3 full weights (July 27, 2026). Comprehensive analysis of the 2.8T-parameter model: KDA architecture, 896-expert MoE, native multimodality, frontier coding benchmarks, and what the open-weight release means for the ecosystem.
Alibaba previews Qwen3.8-Max on July 19, 2026 at WAIC Shanghai — a 2.4T-parameter multimodal MoE claiming 'second only to Fable 5' performance. No benchmarks, no model card, no active-parameter count, no license yet. Open weights promised 'soon.' Available now via Token Plan, Qoder, and QoderWork at 10% preview pricing. Analysis of what's confirmed, what's claimed, and what to wait for.
Thinking Machines Lab releases Inkling on July 15, 2026 — a 975B-parameter open-weights multimodal MoE (41B active) with native text/image/audio, controllable thinking effort, self-improvement via Tinker, and Apache 2.0 licensing. Scores 77.6% on SWE-Bench Verified, 91.4% on VoiceBench, and 73.5% on MMMU Pro, with Inkling-Small (12B active) matching or beating the flagship on key benchmarks.
One major research article published: comprehensive deep-dive on Gemini 3.5 Flash as the agentic frontier model with multimodal reasoning, 1M context, and aggressive pricing. Updated frontier-models wiki page with new benchmark data and enterprise deployment details.
Google DeepMind's Gemini 3.5 Flash, now the default model across Gemini App and AI Mode, ranks #5 in Agentic on BenchLM with 94/100, delivers 76.2% on Terminal-Bench 2.1, and achieves a 68% improvement in token efficiency over Gemini 3 Flash — all at $1.50/$9 per million tokens. With 1M context, 64K output, controllable thinking levels, and native multimodal reasoning, it represents Google's most aggressive price-performance play in the agent-centric era.
Alibaba's Qwen3.7 family — Max (closed-weight flagship, 1M context, SWE-Bench Pro 60.6%, $2.50/$7.50) and Plus (multimodal agent, vision+video, $0.32/$1.28) — represents a strategic pivot from open-weight leadership to closed-weight enterprise competition. Max scores 56.6 on the AA Intelligence Index (#5 overall, highest Chinese model), leads Opus 4.6 on agentic coding benchmarks, and completed a 35-hour autonomous kernel-optimization demo. Plus adds vision-language capabilities at roughly 1/6 the cost. This article analyses the full Qwen3.7 landscape, the open-to-closed pivot, benchmark reality, the verbosity cost trap, and where both models fit in the 2026 frontier.
Google DeepMind's Gemini 3.5 release (May–June 2026) is not a single model — it's a full ecosystem: Flash (frontier coding at Flash-tier pricing), Pro (2M context, Deep Think, enterprise preview), and Audio Live Translate (70+ language real-time speech translation). Combined with the Antigravity platform consolidation and the Gemini CLI retirement on June 18, this is Google's most ambitious AI platform shift since Gemini 1.0. This article maps the full Gemini 3.5 landscape, benchmarks, pricing, and the urgent migration path for developers.
June 5: One new research article — Gemma 4 12B, the encoder-free multimodal laptop model that changes the game. Google DeepMind's 12B dense model eliminates separate vision/audio encoders entirely, runs on 16GB laptops under Apache 2.0, and delivers 78.8% GPQA Diamond. The efficiency revolution now has a multimodal face.
Google DeepMind releases Gemma 4 12B — a 12B dense model with encoder-free multimodal architecture, native audio support, and 256K context. Runs on 16GB laptops under Apache 2.0. Benchmarks approach the 26B MoE sibling at less than half the memory. The most practical multimodal model for local deployment yet.
Google DeepMind released Gemini 3.5 Flash on May 19, 2026 at Google I/O. Built on the Gemini 3 Flash reasoning foundation with thinking levels, it delivers frontier-level agentic and coding performance at 4x the output speed of comparable models. Key results: 76.2% Terminal-Bench 2.1 (beating Gemini 3.1 Pro), 83.6% MCP Atlas, 1656 Elo GDPval-AA, 84.2% CharXiv Reasoning. Priced at $1.50/$9 per 1M tokens with 1M context window. Available via Google Antigravity, Gemini API, Gemini Enterprise Agent Platform, and the Gemini app globally.
Alibaba releases Qwen3.6-35B-A3B, the next iteration of open-source frontier models. Built on community feedback, Qwen3.6 emphasizes agentic coding (frontend workflows, repository-level reasoning), thinking preservation (retaining reasoning context across messages), and refined sparse MoE architecture (40 layers, hybrid Gated DeltaNet + Attention + MoE design). Benchmarks show significant gains over Qwen3.5-35B-A3B and competitive parity with proprietary models.
Comprehensive analysis of Google's Gemma 4 model family—architecture, capabilities, benchmarks, and implications for autonomous agents and on-device AI.
A professional assessment of frontier AI capabilities across text, speech, image, video, and multimodal domains as of March 2026, with performance metrics and source references.