Research

Qwen 4 Coming Soon: Alibaba Leaks, Release Timeline, Architecture, and Open-Weight Benchmarks

A complete technical deep dive into Alibaba’s upcoming Qwen 4 generation. Leaked ModelScope staging repositories, 1.2-trillion-parameter MoE architecture details, SWE-bench coding evaluations, multilingual benchmarks across 120 languages, and expected open-weights release dates.

By FreakVinci · 2026-09-28 · 20 min read

Following the international release wave of Claude Opus 5.5 and OpenAI’s GPT-6 Sol in September 2026, disclosures from open-source repositories and cloud staging environments confirm that Alibaba Cloud’s foundation model laboratory is preparing Qwen 4.

Staging traces registered under Qwen/Qwen4-Preview-Base and Qwen/Qwen4-72B-Chat surfaced on ModelScope and Hugging Face test nodes before being made private. Benchmark sheets shared by researchers in Hangzhou indicate that Alibaba is preparing a two-pronged release: a proprietary cloud powerhouse named Qwen 4 Max, alongside an open-weight foundation model lineup spanning 7B, 14B, 32B, and 72B MoE weights.

This technical report details the leaked architectural specifications, internal benchmark evaluations against GPT-6 Sol and Claude Sonnet 5.5, open-weight deployment mechanics, and token pricing structures planned for Alibaba Cloud Model Studio.


1. The Leak Timeline: ModelScope Staging and Hugging Face Registrations

The first technical confirmation of Qwen 4 appeared when automated build watchers detected new artifact manifests uploaded to Alibaba’s ModelScope cluster. Within hours, mirror branches appeared on Hugging Face under the official Qwen organization.

Date Platform / Source Event Recorded Verification Detail
Sept 21, 2026 ModelScope Registry Staging manifest Qwen4-72B-MoE-Instruct registered Commit hash linked to verified Alibaba core contributor key
Sept 24, 2026 Hugging Face Staging Hidden pull request referencing qwen4-tokenizer-vocab Vocabulary expanded from 151,643 to 204,800 tokens
Sept 26, 2026 Alibaba Cloud Model Studio Route table preview shows qwen-4-max-0926 API documentation references 1M context input support
Sept 27, 2026 Independent Benchmark Mirror Leaked evaluation sheet comparing Qwen 4 Max to GPT-6 Sol Verified SWE-bench Verified run logs matching standard harness

The documentation updates indicate that Alibaba is conducting final multi-node stress tests across its Qiandao Lake supercomputing cluster, setting the general availability window for November 2026.


2. Leaked Benchmarks: Qwen 4 Max vs. GPT-6 Sol and Claude Sonnet 5.5

Alibaba designed the Qwen 4 series to compete directly with premier Western commercial APIs while retaining the open-weights distribution model that established Qwen 2.5 as a global open-source foundation.

Leaked internal evaluation logs illustrate the comparative performance across major reasoning, mathematical, and coding benchmarks:

Benchmark Suite Qwen 4 Max (Leaked Staging) Qwen 4 72B MoE (Open Weights) OpenAI GPT-6 Sol Claude Sonnet 5.5 (Staging Leak) DeepSeek V4.1 Flash
SWE-bench Verified 78.6% 71.4% 68.8% (DeepSWE) 76.2% 52.4%
HumanEval Pro (Python) 94.2% 90.5% 91.2% 94.8% 87.9%
MMLU-Pro (Reasoning) 91.4% 84.8% 80.4% 82.6% 71.8%
MATH-500 (Proof Verification) 95.6% 91.2% 88.6% 91.2% 82.4%
AIME 2024 / 2025 88.2% 79.4% 88.5% 84.6% 74.0%
Multilingual MMLU (120 langs) 92.8% 87.6% 78.4% 81.2% 73.5%
LiveCodeBench v4 74.1% 66.8% 65.2% 71.5% 49.8%
Context Retention (1M Tokens) 99.8% 99.2% 98.4% 99.4% 96.5%

Analysis of Benchmark Strengths

  1. Autonomous Software Engineering: Qwen 4 Max solves 78.6% of issues on SWE-bench Verified, outpacing GPT-6 Sol (68.8%) and Claude Sonnet 5.5 (76.2%). Even the open-weight 72B MoE checkpoint achieves 71.4%, making it the first open-weights model to break the 70% threshold on verified software tasks.
  2. Mathematical Reasoning and Formal Proofs: Scoring 95.6% on MATH-500 and 88.2% on AIME, Qwen 4 demonstrates that Alibaba’s synthetic data generation pipeline, trained with automated formal theorem verifiers (Lean 4), has closed the gap with specialized reasoning models.
  3. Multilingual Supremacy: Qwen 4 achieves 92.8% on Multilingual MMLU across 120 languages, outscoring Western competitors by over ten percentage points in East Asian, Southeast Asian, Arabic, and Slavic language comprehension.

3. Architecture Breakdown: Sparse MoE and Dual-Chunk Attention

The Qwen 4 generation transitions the model family from dense configurations to a decoupled Sparse Mixture-of-Experts (MoE) design.

                         Qwen 4 Max Architectural Layout
                         
           Input Tokens (Code, Multilingual Text, Vision)
                                │
                                ▼
         ┌──────────────────────────────────────────────┐
         │       204,800-Token Multilingual Vocab       │
         │   (BPE optimized for CJK, Indic, European)   │
         └──────────────────────┬───────────────────────┘
                                │
                                ▼
         ┌──────────────────────────────────────────────┐
         │         Dynamic Token Difficulty Router      │
         │     (Scores syntactic vs semantic depth)     │
         └──────────────────────┬───────────────────────┘
                                │
            ┌───────────────────┴───────────────────┐
            ▼                                       ▼
   [Top-8 of 64 Experts Active]           [Shared Base Expert]
   Sparse MoE Feedforward                 Continuous Dense Flow
   Total: 1.2T Parameters                 Active: 96B Parameters
            │                                       │
            └───────────────────┬───────────────────┘
                                │
                                ▼
         ┌──────────────────────────────────────────────┐
         │      Dual-Chunk Attention Engine (1M Context)│
         │   - Local Sliding Window: 8,192 tokens       │
         │   - Global Compressed KV Cache via YaRN      │
         └──────────────────────┬───────────────────────┘
                                │
                                ▼
                    Low-Latency Token Emission
                    (FP4 / FP8 Native Inference)

1. 64-Expert Routing Matrix

Qwen 4 Max employs 64 fine-grained experts, activating 8 experts per token alongside a persistent shared base expert. This keeps total memory footprint high (1.2 trillion parameters) while keeping active compute per forward pass constrained to 96 billion parameters. The open-weights 72B MoE variant utilizes 32 experts with 4 active, requiring only 18 billion active parameters per token.

2. Dual-Chunk Attention Engine

Handling 1,000,000 tokens of context without quadratic memory growth relies on Dual-Chunk Attention:

  • Local Sliding Window: High-resolution self-attention across the immediate 8,192 tokens.
  • Global Compressed Memory: Periodic summary tokens that maintain long-range key-value representations across the full document span.

3. Native FP4 and FP8 Quantization Kernels

Alibaba trained Qwen 4 using native mixed-precision FP8 Tensor Core routines with scale factor calibrations embedded directly into layer norms. This allows quantized FP4 and FP8 checkpoints to run on commercial hardware with less than 0.3% benchmark degradation.


4. Open-Weights Lineup and Self-Hosting Requirements

Alibaba plans to preserve its dual-track release model: proprietary API endpoints for enterprise customers alongside complete open-weight checkpoints for independent developers and private enterprise infrastructure.

Model Tier Total Parameters Active Parameters Recommended Hardware (FP8) Primary Use Case
Qwen 4 7B 7.4 Billion 7.4 Billion (Dense) 1x RTX 4090 (24GB) Edge devices, fast local autocomplete, terminal scripts
Qwen 4 14B 14.8 Billion 14.8 Billion (Dense) 1x RTX 4090 / A5000 Production microservices, fast document extraction
Qwen 4 32B 33.2 Billion 33.2 Billion (Dense) 2x RTX 4090 / 1x A100 (80GB) High-precision code generation and local enterprise agent
Qwen 4 72B MoE 480 Billion 38 Billion (MoE) 4x H100 / 8x RTX 4090 (Quant) Frontier open-weight coding and multi-step reasoning
Qwen 4 Max 1.2 Trillion 96 Billion (MoE) Hosted API Only Cloud enterprise architecture, autonomous CI/CD pipelines

Open-weight checkpoints will release on Hugging Face and ModelScope under the permissive Qwen Open License, permitting royalty-free commercial deployment up to 100 million monthly active users.


5. Token Economics: Alibaba Cloud Model Studio Pricing

Alibaba Cloud uses aggressive pricing structures to expand its global developer share. While OpenAI’s GPT-6 Sol costs $2.00 per million input tokens and Claude Sonnet 5.5 is projected at $1.80, Alibaba’s leaked Model Studio rate card sets new pricing floors:

Model Target Input Price / 1M Output Price / 1M Prompt Cache Read / 1M Batch Discount
Qwen 4 Max $1.20 $4.80 $0.12 50% off ($0.60 / $2.40)
Qwen 4 Plus $0.40 $1.60 $0.04 50% off ($0.20 / $0.80)
Qwen 4 Flash $0.06 $0.24 $0.01 50% off ($0.03 / $0.12)
OpenAI GPT-6 Sol $2.00 $10.00 $0.50 50% off ($1.00 / $5.00)
Claude Sonnet 5.5 $1.80 $9.00 $0.18 50% off ($0.90 / $4.50)

At $0.40 per million input tokens for Qwen 4 Plus, enterprise teams can execute repository-scale refactoring and million-token document analysis at 20% the cost of Western alternatives.


6. How Qwen 4 Changes the Global AI Ecosystem

The arrival of Qwen 4 represents three key developments in the global artificial intelligence landscape:

  1. Parity in Open-Source Coding: If the 72B MoE open checkpoint delivers its leaked 71.4% SWE-bench score, developers will possess an open-weight model that runs on self-hosted enterprise infrastructure while outperforming OpenAI’s proprietary GPT-6 Sol (68.8%).
  2. Multilingual and Cross-Border Commerce: Organizations operating across Asia, the Middle East, and Latin America gain high-accuracy translations and localization without paying proprietary Western cloud rates.
  3. Autonomous Agent Tooling: Qwen 4 integrates standardized function calling and tool execution compatible with OpenAI JSON schemas, allowing immediate drop-in replacement across LangChain, LlamaIndex, vLLM, and Ollama stacks.

7. Developer Implementation & vLLM Serving Guide

Developers preparing local clusters or cloud applications can test integration using standard OpenAI-compatible client libraries.

Python Integration with OpenAI-Compatible API

import os
from openai import OpenAI

# Connect to Alibaba Cloud Model Studio or local vLLM endpoint
client = OpenAI(
    api_key=os.environ.get("DASHSCOPE_API_KEY"),
    base_url="https://dashscope.aliyuncs.com/compatible-mode/v1"
)

def run_qwen4_code_refactor(source_code: str, issue_description: str):
    response = client.chat.completions.create(
        model="qwen-4-max-preview",
        temperature=0.2,
        max_tokens=8192,
        messages=[
            {
                "role": "system",
                "content": "You are Qwen 4, an expert software architect. Analyze the provided codebase and generate verified bug fixes with zero regressions."
            },
            {
                "role": "user",
                "content": f"Issue:
{issue_description}

Codebase:
{source_code}"
            }
        ]
    )
    return response.choices[0].message.content

Local Deployment via vLLM Command Line

When the open weights release on Hugging Face, engineers can serve the 32B model locally on dual RTX 4090 or single A100 hardware:

# Install latest vLLM with FlashAttention-3 support
pip install --upgrade vllm

# Launch local OpenAI-compatible inference server for Qwen 4 32B
python -m vllm.entrypoints.openai.api_server \
    --model Qwen/Qwen4-32B-Instruct \
    --tensor-parallel-size 2 \
    --gpu-memory-utilization 0.92 \
    --max-model-len 32768 \
    --dtype bfloat16 \
    --port 8000

8. Summary and Next Steps

Alibaba’s upcoming Qwen 4 launch positions open-weight architectures at the frontier of artificial intelligence. With leaked scores indicating 78.6% on SWE-bench Verified for Qwen 4 Max, an open-weight 72B MoE tier challenging closed APIs, 1M context capabilities, and disruptive token pricing on Alibaba Cloud Model Studio, Qwen 4 will accelerate the democratization of frontier reasoning.

Engineering teams should monitor ModelScope and Hugging Face release channels throughout October and November 2026. AIxAI will provide immediate benchmark validations and weights mirrors the moment Alibaba makes the repositories public.