3 entries with this tag
Alibaba's Qwen3.7 family — Max (closed-weight flagship, 1M context, SWE-Bench Pro 60.6%, $2.50/$7.50) and Plus (multimodal agent, vision+video, $0.32/$1.28) — represents a strategic pivot from open-weight leadership to closed-weight enterprise competition. Max scores 56.6 on the AA Intelligence Index (#5 overall, highest Chinese model), leads Opus 4.6 on agentic coding benchmarks, and completed a 35-hour autonomous kernel-optimization demo. Plus adds vision-language capabilities at roughly 1/6 the cost. This article analyses the full Qwen3.7 landscape, the open-to-closed pivot, benchmark reality, the verbosity cost trap, and where both models fit in the 2026 frontier.
Moonshot AI released Kimi K2.7 Code on June 12, 2026 — a coding-specialised 1T-parameter MoE model with forced preserve-thinking, ~30% fewer reasoning tokens than K2.6, and strong gains on MCP tool-use benchmarks. This article analyses the architecture, benchmark landscape, pricing, and where K2.7 Code fits in the 2026 agentic coding stack.
Google DeepMind released Gemini 3.5 Flash on May 19, 2026 at Google I/O. Built on the Gemini 3 Flash reasoning foundation with thinking levels, it delivers frontier-level agentic and coding performance at 4x the output speed of comparable models. Key results: 76.2% Terminal-Bench 2.1 (beating Gemini 3.1 Pro), 83.6% MCP Atlas, 1656 Elo GDPval-AA, 84.2% CharXiv Reasoning. Priced at $1.50/$9 per 1M tokens with 1M context window. Available via Google Antigravity, Gemini API, Gemini Enterprise Agent Platform, and the Gemini app globally.