Tools & Products

MiniMax Ships M3.1-Flash-Preview: Coding-Only Architecture, 73.8% SWE-bench, MiniMax Code Live Deployment, and API Economics

MiniMax quietly released M3.1-Flash-Preview across MiniMax Code and platform APIs. Scoring 73.8% on SWE-bench Verified at 165 tokens per second and $0.10/M input pricing, the coding-specialist model enters direct competition with DeepSeek V4.1 Flash and GPT-6 Luna. Full benchmarks, system card breakdown, and API integration.

By FreakVinci · 2026-09-28 · 20 min read

On September 27, 2026, Shanghai-based AI lab MiniMax deployed M3.1-Flash-Preview, a foundation model engineered exclusively for software engineering. As reported by Pandaily and verified across Reddit r/opencode and Startup Fortune, the release occurred without prior teaser campaigns. MiniMax made the model available inside its desktop IDE, MiniMax Code, and enabled API endpoints on platform.minimax.io.

Discussion on X trending topics focused on the model's price-to-performance ratio: M3.1-Flash-Preview scores 73.8% on SWE-bench Verified while generating tokens at 165 tokens per second at an API price point of $0.10 per million input tokens.

This technical report reviews the architectural specifications, empirical benchmark performance against GPT-6 Sol and Claude Sonnet 5.5, API integration steps, and token economics of MiniMax M3.1-Flash-Preview.


1. Release Timeline and Staging Traceability

The deployment of M3.1-Flash-Preview follows MiniMax's transition from general multimodal conversational agents toward focused developer tooling.

Timestamp Source Event / Observed Activity Technical Details
Sept 26, 2026 14:00 UTC MiniMax Platform API API endpoint minimax-m3.1-flash-preview added to route catalog Header: x-minimax-model: m3.1-flash-code
Sept 26, 2026 18:30 UTC r/opencode Community reports M3.1 live in MiniMax Code extension Default model toggle enabled for 1M context codebases
Sept 27, 2026 02:15 UTC Pandaily Official reporting confirms coding-specialist architecture Dual-engine token routing with 165 tps speed verified
Sept 27, 2026 06:40 UTC Startup Fortune Market analysis of Chinese coding models competing globally Cost comparison against OpenAI and Anthropic mid-tiers
Sept 27, 2026 11:00 UTC APIMaster.ai API documentation and latency benchmarks published TTFT verified at 145ms across Asian and US West gateways

The model enters a crowded field alongside DeepSeek V4.1 Flash, OpenAI GPT-6 Luna, and Alibaba Qwen 4 72B MoE. MiniMax differentiated M3.1 by eliminating general conversational chat weights entirely, training solely on syntax trees, test suites, terminal execution traces, and multi-file Git diff histories.


2. Empirical Benchmarks: M3.1-Flash-Preview vs. The Field

MiniMax published initial evaluation results alongside independent testing conducted by open-source agent developers on GitHub.

Benchmark MiniMax M3.1-Flash-Preview OpenAI GPT-6 Sol DeepSeek V4.1 Flash Claude Sonnet 5.5 (Leaked) GPT-6 Luna
SWE-bench Verified 73.8% 68.8% 67.2% 76.2% 48.2%
Terminal-Bench 4.0 76.5% 71.0% 68.4% 79.4% 54.0%
HumanEval Pro (Python) 92.4% 91.2% 89.6% 94.8% 82.5%
MultiPL-E (8 Languages) 88.6% 87.4% 85.1% 91.0% 76.8%
CursorBench 4.0 81.2% 76.5% 74.0% 84.3% 61.2%
Output Generation Speed 165 tps 78 tps 168 tps 115 tps 157 tps
Time-to-First-Token (TTFT) 145 ms 420 ms 160 ms 210 ms 190 ms
Input Price / 1M Tokens $0.10 $2.00 $0.055 $1.80 $0.10
Output Price / 1M Tokens $0.40 $10.00 $0.22 $9.00 $0.50

Benchmark Analysis

  1. Bug Resolution on SWE-bench Verified: M3.1-Flash-Preview scored 73.8%, resolving 369 out of 500 validated GitHub issues without human intervention. This places it 5.0 percentage points ahead of GPT-6 Sol and 6.6 percentage points ahead of DeepSeek V4.1 Flash.
  2. Terminal and Shell Execution: On Terminal-Bench 4.0, which grades autonomous file navigation, bash scripts, and dependency updates, M3.1 achieved 76.5%. It produced fewer syntax loop errors during complex environment setups.
  3. Inference Latency in IDEs: Scoring 165 tokens per second with 145ms TTFT, M3.1 delivers instant inline code completions in MiniMax Code, feeling significantly more responsive than heavy 80-tps models.

3. Architecture Teardown: How M3.1-Flash Achieves High Velocity

MiniMax departed from the standard dense transformer design by implementing a specialized sparse Mixture-of-Experts (MoE) configuration combined with linear attention pre-filters.

                  MiniMax M3.1-Flash Architecture Pipeline
                  
   Raw Codebase Context (Up to 1,000,000 Tokens)
                         │
                         ▼
  ┌────────────────────────────────────────────────────────────┐
  │        Linear Attention Pre-Filter (Prefix Scorer)         │
  │    Compresses static AST structures and header boilerplate │
  └──────────────────────────────┬─────────────────────────────┘
                                 │
                                 ▼
  ┌────────────────────────────────────────────────────────────┐
  │         Sparse MoE Coding Backbone (16 Active Experts)     │
  │   - Specialized Grammar Experts (TypeScript, Rust, Go)     │
  │   - Diff & Patch Expert (Unified Diff format compliance)   │
  │   - Shell & Terminal Execution Expert                      │
  └──────────────────────────────┬─────────────────────────────┘
                                 │
                                 ▼
  ┌────────────────────────────────────────────────────────────┐
  │         Speculative Token Drafter (FP8 Kernel Cluster)     │
  │   Generates 4 candidate tokens per forward pass (~165 tps) │
  └──────────────────────────────┬─────────────────────────────┘
                                 │
                                 ▼
                 Syntactically Valid Code / Git Patch

1. Abstract Syntax Tree (AST) Alignment Pre-Training

MiniMax filtered its pre-training data to exclude unstructured web crawl text, training predominantly on 14 trillion tokens of syntax-parsed source code, unit test pairs, issue-resolution pairs, and compiler diagnostic logs. This targeted data mixture reduces generation hallucinations on type signatures and imported package paths.

2. Fast Speculative Token Drafter

By pairing the main MoE transformer with a dedicated 3-billion-parameter speculative drafter running in 8-bit floating point (FP8), M3.1 generates 4 candidate tokens in parallel, verifying them against the primary model in a single memory cycle.

3. 1,000,000 Token Linear Attention Cache

The model maintains a 1M-token context buffer. Static files (such as documentation headers, package manifests, and unchanged libraries) pass through a linear attention projection that compresses the KV cache footprint by 65%, keeping memory consumption within manageable hardware thresholds during large refactoring runs.


4. Token Economics: Pricing Comparison

MiniMax structured its API rates to compete directly with low-cost inference providers while undercutting western frontier platforms by more than an order of magnitude.

Model Input Price / 1M Prompt Cache Read / 1M Output Price / 1M Batch Job Input / 1M Batch Job Output / 1M
MiniMax M3.1-Flash-Preview $0.10 $0.02 $0.40 $0.05 $0.20
DeepSeek V4.1 Flash $0.055 $0.014 $0.22 $0.0275 $0.11
OpenAI GPT-6 Luna $0.10 $0.025 $0.50 $0.05 $0.25
OpenAI GPT-6 Sol $2.00 $0.50 $10.00 $1.00 $5.00
Claude Sonnet 5.5 (Expected) $1.80 $0.18 $9.00 $0.90 $4.50

Cost Simulation: 100 Million Input Tokens (75% Cache Hit) + 20 Million Output Tokens

Evaluating a mid-sized software team running continuous automated pull request reviews:

  1. Claude Sonnet 5.5:

    • Cached Input: 75M × $0.18 = $13.50
    • Uncached Input: 25M × $1.80 = $45.00
    • Output: 20M × $9.00 = $180.00
    • Total Monthly Cost: $238.50
  2. OpenAI GPT-6 Sol:

    • Cached Input: 75M × $0.50 = $37.50
    • Uncached Input: 25M × $2.00 = $50.00
    • Output: 20M × $10.00 = $200.00
    • Total Monthly Cost: $287.50
  3. MiniMax M3.1-Flash-Preview:

    • Cached Input: 75M × $0.02 = $1.50
    • Uncached Input: 25M × $0.10 = $2.50
    • Output: 20M × $0.40 = $8.00
    • Total Monthly Cost: $12.00 (95% reduction compared to Sonnet 5.5)

For continuous software verification pipelines, M3.1 delivers strong coding accuracy at a fraction of typical frontier operating expenses.


5. MiniMax Code: Native IDE Integration

MiniMax shipped M3.1 directly into MiniMax Code, an Electron-based development fork engineered around autonomous coding agents.

Core IDE Capabilities

  1. Repository-Wide Semantic Indexing: MiniMax Code indexes symbols, type declarations, and dependency trees into a local SQLite vector database, injecting relevant context into the 1M token buffer automatically.
  2. Deterministic Git Diff Generation: Rather than re-emitting whole files, M3.1 outputs unified diffs with line ranges, preventing accidental deletions in large source files.
  3. Integrated Terminal Execution: When running test suites, MiniMax Code captures failed stack traces and feeds them back to M3.1 for automated debugging loops.

6. API Integration Guide

The MiniMax API adheres to the OpenAI-compatible chat completions interface, simplifying integration with existing client libraries and agent frameworks like Aider, Cursor, and OpenCode.

TypeScript / Node.js Implementation

import OpenAI from 'openai';

// Initialize MiniMax API client
const minimax = new OpenAI({
  apiKey: process.env.MINIMAX_API_KEY,
  baseURL: 'https://api.minimax.io/v1',
});

async function runMiniMaxCodeRefactor(codebaseFiles: string, issuePrompt: string) {
  const response = await minimax.chat.completions.create({
    model: 'minimax-m3.1-flash-preview',
    messages: [
      {
        role: 'system',
        content: 'You are an autonomous senior software engineer. Output changes only in verified unified diff format.',
      },
      {
        role: 'user',
        content: `Repository Context:\n${codebaseFiles}\n\nTask:\n${issuePrompt}`,
      },
    ],
    temperature: 0.1,
    max_tokens: 16384,
  });

  const diffOutput = response.choices[0].message.content;
  console.log('Generated Patch:\n', diffOutput);
  console.log('Usage metrics:', response.usage);
}

Python Implementation with Automated Test Execution

import os
import subprocess
from openai import OpenAI

client = OpenAI(
    api_key=os.environ.get("MINIMAX_API_KEY"),
    base_url="https://api.minimax.io/v1"
)

def solve_github_issue(issue_description: str, repo_path: str):
    # Read relevant files
    prompt = f"Solve the following issue in {repo_path}:\n{issue_description}"
    
    response = client.chat.completions.create(
        model="minimax-m3.1-flash-preview",
        messages=[
            {"role": "system", "content": "You are a coding specialist. Return strict bash execution commands to patch and test."},
            {"role": "user", "content": prompt}
        ],
        temperature=0.0,
        max_tokens=8192
    )
    
    script = response.choices[0].message.content
    print("Executing fix script...")
    return script

7. Strategic Outlook: The Specialization of Developer Models

The unannounced launch of MiniMax M3.1-Flash-Preview highlights a clear trend in AI development: domain specialization. Instead of scaling multi-trillion-parameter models across all tasks, labs are building domain-specific architectures trained exclusively on coding and terminal execution.

With 73.8% on SWE-bench Verified, 165 tokens per second throughput, and $0.10/M input pricing, MiniMax M3.1-Flash-Preview provides engineering teams with a high-speed, cost-effective coding engine for continuous integration and interactive IDE workflows.