QWEN3.7-MAX

Qwen3.7-Max: Alibaba's frontier agent model

Released May 20, 2026. GPQA Diamond 92.4. Sustains 35+ hours of autonomous execution. 1M-token context, native web search. $1.65/$4.95 per million tokens.

Try Qwen3.7-Max on Council See it in a council

Qwen3.7-Max is Alibaba's frontier agent model, released May 20, 2026, and the strongest entry yet in the Qwen Max line. Its two headline numbers define it: 92.4 on GPQA Diamond — frontier-class scientific reasoning — and sustained autonomous execution beyond 35 hours, the endurance metric Alibaba optimized for and the longest claimed in Council's lineup. It ships a 1M-token context window and native web search via DashScope's enable_search; it is text-only, with no vision input. Pricing on DashScope global is $1.65 per million input tokens and $4.95 per million output — frontier reasoning at well under half the output price of the Western flagships. Versus its predecessor Qwen3.6-Max (April 20, 2026; $1.04/$6.24, 262K context, #1 on six coding benchmarks at launch), 3.7-Max pivots from benchmark coding wins to agent endurance: nearly 4× the context, cheaper output, and a training focus on staying coherent over very long autonomous runs. The case for Qwen3.7-Max: marathon agent workloads, scientific reasoning per dollar, and multilingual work. The case against: anything that needs vision, and refactor-grade patch quality where Claude leads.

Specs

ProviderAlibaba (Qwen, via DashScope)
ReleasedMay 20, 2026
Context window1,000,000 tokens
GPQA Diamond92.4
Autonomous executionSustains 35+ hours
VisionNo — text only
Web searchNative (DashScope enable_search)
Price (in / out per 1M)$1.65 / $4.95 (DashScope global)

Best at

Where it loses

Qwen3.7-Max vs Qwen3.6-Max

Qwen3.6-Max (April 20, 2026) made its name on coding benchmarks — #1 on six of them at launch, including SWE-bench Pro and Terminal-Bench 2.0 — with a 262K context at $1.04/$6.24. One month later, 3.7-Max changed the axis of competition: context grows nearly 4× to 1M, output gets ~20% cheaper ($4.95 vs $6.24), input costs more ($1.65 vs $1.04), and the training target shifts from benchmark wins to agent endurance — GPQA Diamond 92.4 and 35+ hour autonomous runs. Both remain in Council's lineup: 3.6-Max is still a superb pure-coding seat; 3.7-Max is the one you give a long leash.

How Council AI runs it

Qwen3.7-Max is available on every paid Council AI plan (Plus, Pro, Ultra). In a council it answers in parallel with models from up to 8 other labs — its DashScope-native web search keeps it current, and its cost profile means seating it rarely forces dropping another voice. The moderator scores agreement across all answers; Qwen's independent training lineage makes its agreement with Western flagships a strong correctness signal, and its GPQA-class reasoning means its dissents on technical questions deserve a second look rather than a dismissal.

Frequently asked questions

How much does Qwen3.7-Max cost via API?

$1.65 per million input tokens and $4.95 per million output tokens on DashScope global pricing. Council AI bundles it into monthly plan budgets.

What is Qwen3.7-Max's context window?

1,000,000 tokens — nearly 4× its predecessor Qwen3.6-Max's 262K.

What does '35+ hours of autonomous execution' mean?

Alibaba trained and evaluated 3.7-Max on very long unattended agent runs — the model sustains coherent plan-act-observe loops for 35+ hours without human intervention. It's the longest such claim in Council's lineup.

How does Qwen3.7-Max compare to Qwen3.6-Max?

3.6-Max (Apr 2026) won on coding benchmarks with 262K context at $1.04/$6.24. 3.7-Max (May 2026) targets agents: 1M context, GPQA Diamond 92.4, 35+ hour autonomy, $1.65/$4.95. Both are in Council's active lineup.

Does Qwen3.7-Max support vision or web search?

Web search yes — native via DashScope's enable_search. Vision no — it's text-only; pair it with a multimodal seat like Gemini 3.5 Flash in a council.

Is Qwen3.7-Max available in Council AI?

Yes — on all paid plans (Plus, Pro, Ultra), running in parallel councils with models from 8 other labs and consensus scoring on every answer.