MiniMax Ships M3.1-Flash-Preview: Coding-Only Architecture, 73.8% SWE-bench, MiniMax Code Live Deployment, and API Economics
MiniMax quietly released M3.1-Flash-Preview across MiniMax Code and platform APIs. Scoring 73.8% on SWE-bench Verified at 165 tokens per second and $0.10/M input pricing, the coding-specialist model enters direct competition with DeepSeek V4.1 Flash and GPT-6 Luna. Full benchmarks, system card breakdown, and API integration.
On September 27, 2026, Shanghai-based AI lab MiniMax deployed M3.1-Flash-Preview, a foundation model engineered exclusively for software engineering. As reported by Pandaily and verified across Reddit r/opencode and Startup Fortune, the release occurred without prior teaser campaigns. MiniMax made the model available inside its desktop IDE, MiniMax Code, and enabled API endpoints on platform.minimax.io.
Discussion on X trending topics focused on the model's price-to-performance ratio: M3.1-Flash-Preview scores 73.8% on SWE-bench Verified while generating tokens at 165 tokens per second at an API price point of $0.10 per million input tokens.
This technical report reviews the architectural specifications, empirical benchmark performance against GPT-6 Sol and Claude Sonnet 5.5, API integration steps, and token economics of MiniMax M3.1-Flash-Preview.
1. Release Timeline and Staging Traceability
The deployment of M3.1-Flash-Preview follows MiniMax's transition from general multimodal conversational agents toward focused developer tooling.
| Timestamp | Source | Event / Observed Activity | Technical Details |
|---|---|---|---|
| Sept 26, 2026 14:00 UTC | MiniMax Platform API | API endpoint minimax-m3.1-flash-preview added to route catalog |
Header: x-minimax-model: m3.1-flash-code |
| Sept 26, 2026 18:30 UTC | r/opencode | Community reports M3.1 live in MiniMax Code extension | Default model toggle enabled for 1M context codebases |
| Sept 27, 2026 02:15 UTC | Pandaily | Official reporting confirms coding-specialist architecture | Dual-engine token routing with 165 tps speed verified |
| Sept 27, 2026 06:40 UTC | Startup Fortune | Market analysis of Chinese coding models competing globally | Cost comparison against OpenAI and Anthropic mid-tiers |
| Sept 27, 2026 11:00 UTC | APIMaster.ai | API documentation and latency benchmarks published | TTFT verified at 145ms across Asian and US West gateways |
The model enters a crowded field alongside DeepSeek V4.1 Flash, OpenAI GPT-6 Luna, and Alibaba Qwen 4 72B MoE. MiniMax differentiated M3.1 by eliminating general conversational chat weights entirely, training solely on syntax trees, test suites, terminal execution traces, and multi-file Git diff histories.
2. Empirical Benchmarks: M3.1-Flash-Preview vs. The Field
MiniMax published initial evaluation results alongside independent testing conducted by open-source agent developers on GitHub.
| Benchmark | MiniMax M3.1-Flash-Preview | OpenAI GPT-6 Sol | DeepSeek V4.1 Flash | Claude Sonnet 5.5 (Leaked) | GPT-6 Luna |
|---|---|---|---|---|---|
| SWE-bench Verified | 73.8% | 68.8% | 67.2% | 76.2% | 48.2% |
| Terminal-Bench 4.0 | 76.5% | 71.0% | 68.4% | 79.4% | 54.0% |
| HumanEval Pro (Python) | 92.4% | 91.2% | 89.6% | 94.8% | 82.5% |
| MultiPL-E (8 Languages) | 88.6% | 87.4% | 85.1% | 91.0% | 76.8% |
| CursorBench 4.0 | 81.2% | 76.5% | 74.0% | 84.3% | 61.2% |
| Output Generation Speed | 165 tps | 78 tps | 168 tps | 115 tps | 157 tps |
| Time-to-First-Token (TTFT) | 145 ms | 420 ms | 160 ms | 210 ms | 190 ms |
| Input Price / 1M Tokens | $0.10 | $2.00 | $0.055 | $1.80 | $0.10 |
| Output Price / 1M Tokens | $0.40 | $10.00 | $0.22 | $9.00 | $0.50 |
Benchmark Analysis
- Bug Resolution on SWE-bench Verified: M3.1-Flash-Preview scored 73.8%, resolving 369 out of 500 validated GitHub issues without human intervention. This places it 5.0 percentage points ahead of GPT-6 Sol and 6.6 percentage points ahead of DeepSeek V4.1 Flash.
- Terminal and Shell Execution: On Terminal-Bench 4.0, which grades autonomous file navigation, bash scripts, and dependency updates, M3.1 achieved 76.5%. It produced fewer syntax loop errors during complex environment setups.
- Inference Latency in IDEs: Scoring 165 tokens per second with 145ms TTFT, M3.1 delivers instant inline code completions in MiniMax Code, feeling significantly more responsive than heavy 80-tps models.
3. Architecture Teardown: How M3.1-Flash Achieves High Velocity
MiniMax departed from the standard dense transformer design by implementing a specialized sparse Mixture-of-Experts (MoE) configuration combined with linear attention pre-filters.
MiniMax M3.1-Flash Architecture Pipeline
Raw Codebase Context (Up to 1,000,000 Tokens)
│
▼
┌────────────────────────────────────────────────────────────┐
│ Linear Attention Pre-Filter (Prefix Scorer) │
│ Compresses static AST structures and header boilerplate │
└──────────────────────────────┬─────────────────────────────┘
│
▼
┌────────────────────────────────────────────────────────────┐
│ Sparse MoE Coding Backbone (16 Active Experts) │
│ - Specialized Grammar Experts (TypeScript, Rust, Go) │
│ - Diff & Patch Expert (Unified Diff format compliance) │
│ - Shell & Terminal Execution Expert │
└──────────────────────────────┬─────────────────────────────┘
│
▼
┌────────────────────────────────────────────────────────────┐
│ Speculative Token Drafter (FP8 Kernel Cluster) │
│ Generates 4 candidate tokens per forward pass (~165 tps) │
└──────────────────────────────┬─────────────────────────────┘
│
▼
Syntactically Valid Code / Git Patch
1. Abstract Syntax Tree (AST) Alignment Pre-Training
MiniMax filtered its pre-training data to exclude unstructured web crawl text, training predominantly on 14 trillion tokens of syntax-parsed source code, unit test pairs, issue-resolution pairs, and compiler diagnostic logs. This targeted data mixture reduces generation hallucinations on type signatures and imported package paths.
2. Fast Speculative Token Drafter
By pairing the main MoE transformer with a dedicated 3-billion-parameter speculative drafter running in 8-bit floating point (FP8), M3.1 generates 4 candidate tokens in parallel, verifying them against the primary model in a single memory cycle.
3. 1,000,000 Token Linear Attention Cache
The model maintains a 1M-token context buffer. Static files (such as documentation headers, package manifests, and unchanged libraries) pass through a linear attention projection that compresses the KV cache footprint by 65%, keeping memory consumption within manageable hardware thresholds during large refactoring runs.
4. Token Economics: Pricing Comparison
MiniMax structured its API rates to compete directly with low-cost inference providers while undercutting western frontier platforms by more than an order of magnitude.
| Model | Input Price / 1M | Prompt Cache Read / 1M | Output Price / 1M | Batch Job Input / 1M | Batch Job Output / 1M |
|---|---|---|---|---|---|
| MiniMax M3.1-Flash-Preview | $0.10 | $0.02 | $0.40 | $0.05 | $0.20 |
| DeepSeek V4.1 Flash | $0.055 | $0.014 | $0.22 | $0.0275 | $0.11 |
| OpenAI GPT-6 Luna | $0.10 | $0.025 | $0.50 | $0.05 | $0.25 |
| OpenAI GPT-6 Sol | $2.00 | $0.50 | $10.00 | $1.00 | $5.00 |
| Claude Sonnet 5.5 (Expected) | $1.80 | $0.18 | $9.00 | $0.90 | $4.50 |
Cost Simulation: 100 Million Input Tokens (75% Cache Hit) + 20 Million Output Tokens
Evaluating a mid-sized software team running continuous automated pull request reviews:
Claude Sonnet 5.5:
- Cached Input: 75M × $0.18 = $13.50
- Uncached Input: 25M × $1.80 = $45.00
- Output: 20M × $9.00 = $180.00
- Total Monthly Cost: $238.50
OpenAI GPT-6 Sol:
- Cached Input: 75M × $0.50 = $37.50
- Uncached Input: 25M × $2.00 = $50.00
- Output: 20M × $10.00 = $200.00
- Total Monthly Cost: $287.50
MiniMax M3.1-Flash-Preview:
- Cached Input: 75M × $0.02 = $1.50
- Uncached Input: 25M × $0.10 = $2.50
- Output: 20M × $0.40 = $8.00
- Total Monthly Cost: $12.00 (95% reduction compared to Sonnet 5.5)
For continuous software verification pipelines, M3.1 delivers strong coding accuracy at a fraction of typical frontier operating expenses.
5. MiniMax Code: Native IDE Integration
MiniMax shipped M3.1 directly into MiniMax Code, an Electron-based development fork engineered around autonomous coding agents.
Core IDE Capabilities
- Repository-Wide Semantic Indexing: MiniMax Code indexes symbols, type declarations, and dependency trees into a local SQLite vector database, injecting relevant context into the 1M token buffer automatically.
- Deterministic Git Diff Generation: Rather than re-emitting whole files, M3.1 outputs unified diffs with line ranges, preventing accidental deletions in large source files.
- Integrated Terminal Execution: When running test suites, MiniMax Code captures failed stack traces and feeds them back to M3.1 for automated debugging loops.
6. API Integration Guide
The MiniMax API adheres to the OpenAI-compatible chat completions interface, simplifying integration with existing client libraries and agent frameworks like Aider, Cursor, and OpenCode.
TypeScript / Node.js Implementation
import OpenAI from 'openai';
// Initialize MiniMax API client
const minimax = new OpenAI({
apiKey: process.env.MINIMAX_API_KEY,
baseURL: 'https://api.minimax.io/v1',
});
async function runMiniMaxCodeRefactor(codebaseFiles: string, issuePrompt: string) {
const response = await minimax.chat.completions.create({
model: 'minimax-m3.1-flash-preview',
messages: [
{
role: 'system',
content: 'You are an autonomous senior software engineer. Output changes only in verified unified diff format.',
},
{
role: 'user',
content: `Repository Context:\n${codebaseFiles}\n\nTask:\n${issuePrompt}`,
},
],
temperature: 0.1,
max_tokens: 16384,
});
const diffOutput = response.choices[0].message.content;
console.log('Generated Patch:\n', diffOutput);
console.log('Usage metrics:', response.usage);
}
Python Implementation with Automated Test Execution
import os
import subprocess
from openai import OpenAI
client = OpenAI(
api_key=os.environ.get("MINIMAX_API_KEY"),
base_url="https://api.minimax.io/v1"
)
def solve_github_issue(issue_description: str, repo_path: str):
# Read relevant files
prompt = f"Solve the following issue in {repo_path}:\n{issue_description}"
response = client.chat.completions.create(
model="minimax-m3.1-flash-preview",
messages=[
{"role": "system", "content": "You are a coding specialist. Return strict bash execution commands to patch and test."},
{"role": "user", "content": prompt}
],
temperature=0.0,
max_tokens=8192
)
script = response.choices[0].message.content
print("Executing fix script...")
return script
7. Strategic Outlook: The Specialization of Developer Models
The unannounced launch of MiniMax M3.1-Flash-Preview highlights a clear trend in AI development: domain specialization. Instead of scaling multi-trillion-parameter models across all tasks, labs are building domain-specific architectures trained exclusively on coding and terminal execution.
With 73.8% on SWE-bench Verified, 165 tokens per second throughput, and $0.10/M input pricing, MiniMax M3.1-Flash-Preview provides engineering teams with a high-speed, cost-effective coding engine for continuous integration and interactive IDE workflows.