← Back to Home

#long-horizon

3 entries with this tag

🔬 research2026-08-05T00:00:00.000Z

Qwen3.8-Max: 2.4T Parameters, Open Weights, and the First Model to Code Autonomously for 16 Days

Alibaba released Qwen3.8-Max on August 3, 2026 — a 2.4T-parameter sparse MoE model with 95B active parameters, 1M-token context, and open weights coming next week. Covers the architecture, benchmark results (86.6 Terminal-Bench 2.1, 56.6 DeepSWE 1.1, 73.5 FrontierSWE), the 16-day autonomous coding project (oh-my-cli), research paper reproduction with improvement, multimodal capabilities, and the strategic implications for the open-weight frontier.

#qwen#alibaba#moe#open-weight#autonomous-coding#long-horizon#multimodal
🔬 research2026-07-01T00:00:00.000Z

Qwen3.7-Max: The Agent-Centric Era — Long-Horizon Execution, Language World Models, and Alibaba's Frontier Push

Alibaba's Qwen team released Qwen3.7-Max, a next-generation proprietary flagship designed for the agent-centric era with 1M-token context, deep reasoning, and strong coding/agent benchmarks. Paired with the open-source Qwen-AgentWorld-35B-A3B (a language world model covering 7 agent domains), the release positions Qwen3.7-Max at 90/100 on BenchLM overall, #6 in coding, and ahead of DeepSeek-V4-Pro on Terminal-Bench 2.0 — while remaining within 0.2 points on SWE-bench Verified.

#Qwen#Qwen3.7#Alibaba#Agent-Centric#Long-Horizon#Language World Models#Coding
🔬 research2026-05-20T00:00:00.000Z

Qwen3.7-Max: The Agent Frontier — Comparing Alibaba's Latest Proprietary Model Against the April 2026 Tier

Qwen3.7-Max is Alibaba's new proprietary agent foundation model, released May 20, 2026. It challenges the April 2026 frontier trio (DeepSeek-V4-Pro, GPT-5.5, Claude Opus 4.7) by combining coding agent leadership (69.7% Terminal-Bench, 60.6% SWE-Pro), office productivity (87% SpreadsheetBench), and 35-hour autonomous execution. Available via Alibaba Cloud Model Studio API only.

#qwen3.7#frontier#agents#comparison#alibaba#coding-agent#long-horizon