Loading...
17 entries with this tag
On August 14, 2026, Alibaba's Qwen team released Qwen3.8-27B β a 27B dense, native vision-language model with hybrid Gated DeltaNet + Gated Attention architecture, flexible thinking control, and Apache 2.0 licensing. The model delivers 73.0 on Terminal Bench 2.1 (within 5 points of Opus 4.6 Max), 61.7 on SWE-bench Pro, 84.3 on OSWorld-Verified, and 90.0 on MathVision, all in a model that fits on a single consumer GPU. Covers architecture, text and vision benchmarks, deployment guidance, and strategic implications for the local AI landscape.
One new research article published: Qwen3.8-Max, Alibaba's 2.4T-parameter sparse MoE model with open weights coming next week, 16-day autonomous coding project, and the first model to reproduce and improve upon a research paper without human intervention.
Alibaba released Qwen3.8-Max on August 3, 2026 β a 2.4T-parameter sparse MoE model with 95B active parameters, 1M-token context, and open weights coming next week. Covers the architecture, benchmark results (86.6 Terminal-Bench 2.1, 56.6 DeepSWE 1.1, 73.5 FrontierSWE), the 16-day autonomous coding project (oh-my-cli), research paper reproduction with improvement, multimodal capabilities, and the strategic implications for the open-weight frontier.
One new research article published: deep analysis of Alibaba's Qwen3.8-Max-Preview announcement at WAIC Shanghai β a 2.4T-parameter multimodal MoE claiming 'second only to Fable 5' with no benchmarks, no model card, and an open-weight release promised 'soon.'
Alibaba previews Qwen3.8-Max on July 19, 2026 at WAIC Shanghai β a 2.4T-parameter multimodal MoE claiming 'second only to Fable 5' performance. No benchmarks, no model card, no active-parameter count, no license yet. Open weights promised 'soon.' Available now via Token Plan, Qoder, and QoderWork at 10% preview pricing. Analysis of what's confirmed, what's claimed, and what to wait for.
One major research article published: comprehensive analysis of Qwen3.7-Max, Alibaba's agent-centric frontier model, paired with the open-source Qwen-AgentWorld language world models. Updated frontier-models and Qwen entity wiki pages.
Alibaba's Qwen team released Qwen3.7-Max, a next-generation proprietary flagship designed for the agent-centric era with 1M-token context, deep reasoning, and strong coding/agent benchmarks. Paired with the open-source Qwen-AgentWorld-35B-A3B (a language world model covering 7 agent domains), the release positions Qwen3.7-Max at 90/100 on BenchLM overall, #6 in coding, and ahead of DeepSeek-V4-Pro on Terminal-Bench 2.0 β while remaining within 0.2 points on SWE-bench Verified.
Alibaba's Tongyi Lab released the Qwen-Robot Suite on June 16, 2026 β three foundation models (Qwen-RobotNav, Qwen-RobotManip, Qwen-RobotWorld) that bridge the gap between digital intelligence and physical action. RobotManip tops RoboChallenge with 20% relative improvement over Ο0.5, RobotNav achieves 76.5% on VLN-CE RxR, and RobotWorld ranks 1st on EWMBench. All models are open-weight with technical reports on arXiv.
Alibaba's Qwen3.7 family β Max (closed-weight flagship, 1M context, SWE-Bench Pro 60.6%, $2.50/$7.50) and Plus (multimodal agent, vision+video, $0.32/$1.28) β represents a strategic pivot from open-weight leadership to closed-weight enterprise competition. Max scores 56.6 on the AA Intelligence Index (#5 overall, highest Chinese model), leads Opus 4.6 on agentic coding benchmarks, and completed a 35-hour autonomous kernel-optimization demo. Plus adds vision-language capabilities at roughly 1/6 the cost. This article analyses the full Qwen3.7 landscape, the open-to-closed pivot, benchmark reality, the verbosity cost trap, and where both models fit in the 2026 frontier.
Qwen3.6-27B (April 22, 2026) is a dense 27B open-weight model that outperforms Alibaba's own 397B MoE on agentic coding benchmarks. With 77.2% SWE-Bench Verified, perfect 100/100 tool calling, Thinking Preservation, and Apache 2.0 licensing, it rewrites the open-weight efficiency curve. A 27B model fitting on a single H100 that matches frontier-tier 397B MoE performance proves parameter count is no longer the only quality lever.
May 21: Two major research articles published. Qwen-SEA-LION-v4.5-27B represents Phase 3 of regional specialization β distilling 397B reasoning into 27B for SEA languages. Qwen3.7-Max enters the frontier as an agent-first model leading on SWE-Pro (60.6%) and 35-hour autonomous execution. The frontier is fragmenting into specialized niches.
Comprehensive comparison of three leading open-source models for autonomous agent deployment: Alibaba Qwen3.6-35B-A3B (thinking preservation + efficiency), DeepSeek-V4-Pro (code generation + reasoning), and Google Gemma 4 31B (balanced frontier + multimodal + function-calling). Benchmarks, architecture, and deployment guidance from official sources only.
Qwen3.6-35B-A3B release analysis: Thinking preservation breakthrough, agentic coding leadership (+5-11% improvements), and open-source frontier maturity validated for local deployment.
Alibaba releases Qwen3.6-35B-A3B, the next iteration of open-source frontier models. Built on community feedback, Qwen3.6 emphasizes agentic coding (frontend workflows, repository-level reasoning), thinking preservation (retaining reasoning context across messages), and refined sparse MoE architecture (40 layers, hybrid Gated DeltaNet + Attention + MoE design). Benchmarks show significant gains over Qwen3.5-35B-A3B and competitive parity with proprietary models.
Comprehensive head-to-head comparison of Claude Haiku 4.5 (API, closed), Qwen3.5-4B (open-source), and Gemma 4 E4B (open-source)βthree leading small models for edge deployment, autonomous agents, and cost-optimized inference.
Head-to-head benchmark analysis of Qwen3.5-4B and Gemma 4 E4Bβtwo leading 4B-class models for edge AI, local inference, and autonomous agents.
Alibaba's Qwen model family β open-weight leaders (Qwen3.6-27B) and closed-weight pivot (Qwen3.7 Max/Plus); SEA-LION regional variant