NVIDIA Unveils $4,999 64GB DGX Spark: Personal Grace Blackwell AI Supercomputer for Sub-100B Parameter Models
NVIDIA introduced an accessible 64GB configuration of its DGX Spark personal AI supercomputer priced at $4,999, powered by the GB10 Grace Blackwell Superchip to deliver local, privacy-focused inference for models up to 100B parameters.
NVIDIA expanded its personal computing portfolio by officially introducing a 64GB variant of the DGX Spark desktop AI supercomputer at an entry price of $4,999.
Scheduled for worldwide retail availability on October 23 through hardware partners including Acer, ASUS, Dell, Gigabyte, HP, and MSI, the machine provides individual AI researchers, engineering teams, and privacy-sensitive enterprises with a dedicated local execution environment for models spanning up to 100 billion parameters.
Specifications: 64GB DGX Spark vs. 128GB High-End Edition
The 64GB model retains the same computational die as its larger sibling, reducing high-bandwidth memory density to counter acute industry shortages.
| Technical Metric | DGX Spark (64GB Entry) | DGX Spark (128GB Pro) |
|---|---|---|
| Launch Price | $4,999 | $6,950 (inflated from $3,999 launch target) |
| Superchip Architecture | NVIDIA GB10 Grace Blackwell | NVIDIA GB10 Grace Blackwell |
| CPU Complex | 72-core Arm Neoverse V3 | 72-core Arm Neoverse V3 |
| Unified Memory | 64GB LPDDR5X + Unified HBM | 128GB Unified HBM3e |
| Memory Bandwidth | 850 GB/s aggregate | 1,450 GB/s aggregate |
| FP4 Tensor Core Compute | 1.8 PetaFLOPS | 1.8 PetaFLOPS |
| FP8 Tensor Core Compute | 900 TeraFLOPS | 900 TeraFLOPS |
| Max Single-Node Model Size | 70B-100B params (FP8/FP4) | 140B-200B params (FP8) |
| Dual-Unit NVLink Scaling | Supported (1.7x scaling) | Supported (1.8x scaling) |
| Thermal Design Power (TDP) | 350 Watts (Whisper-quiet fan) | 480 Watts (Dual closed-loop liquid) |
| Availability Date | October 23, 2026 | In market (constrained allocation) |
Resolving the HBM Shortage Bottleneck
During early 2026, severe packaging bottlenecks around TSMC's CoWoS-S substrate and High-Bandwidth Memory (HBM3e) layers pushed production costs upward. The original 128GB DGX Spark saw street prices escalate by nearly 75%, climbing to $6,950 and putting local hardware out of reach for independent developers.
The 64GB configuration addresses this pricing cliff. By utilizing a hybrid memory arrangement that combines 32GB of on-die high-speed cache with 32GB of high-density wide-bus memory, NVIDIA maintains sub-10ms time-to-first-token (TTFT) metrics while cutting the bill of materials substantially.
Local Model Execution Benchmarks
Running frontier open-weights models on the 64GB DGX Spark produces generation speeds comparable to hosted cloud APIs, but with zero data leakage:
Model: Llama-3.3-70B-Instruct (FP8 Quantized)
Context Window: 8,192 tokens
Concurrency: Single User Interactive Stream
• Time-To-First-Token (TTFT): 112 ms
• Sustained Output Speed: 48.6 tokens/sec
• Active Memory Allocated: 41.2 GB / 64 GB
• System Chassis Acoustic Level: 28 dB (whisper quiet)
Multi-Node NVLink Clustering
For workloads exceeding 64GB of memory—such as running unquantized 70B models or medium-sized fine-tuning runs—two DGX Spark units can connect directly through an external NVLink Bridge.
┌────────────────────────┐ ┌────────────────────────┐
│ DGX Spark Node A │ │ DGX Spark Node B │
│ • GB10 Superchip │◄── NVLink Bridge ─►│ • GB10 Superchip │
│ • 64GB Memory │ (400 GB/s) │ • 64GB Memory │
└────────────────────────┘ └────────────────────────┘
▲ ▲
└───────────────── Combined Pool ─────────────┘
• 128GB Unified Space
• 1.7x Aggregate Throughput
This dual-node configuration pools memory into an aggregate 128GB address space, delivering 1.7x throughput gains without requiring complex 400GbE InfiniBand switches or datacenter racking.
Privacy and Compliance Advantages
As organizations confront legal liabilities surrounding the transmission of proprietary IP to third-party cloud LLMs, local workstations serve a strategic operational role.
The $4,999 DGX Spark operates on standard 110V/220V household wall power, produces minimal acoustic noise during full-throttle inference, and provides an air-gapped environment where sensitive clinical trials, legal discovery documents, and internal source code never leave the local office perimeter.