# NVIDIA Unveils 64GB DGX Spark Desktop AI System for Local Agent Workloads
NVIDIA has introduced a new 64GB variant of its DGX Spark desktop AI workstation, built around the GB10 Grace Blackwell superchip. The system is now available through partners including Acer, ASUS, Dell, Gigabyte, HP, and MSI, and is designed to give developers and researchers a fully self-contained machine for running open models, deploying always-on AI agents, and performing fine-tuning — all without relying on cloud APIs or metered token services.
## Why the 64GB Configuration Matters Now
Agent-driven AI workloads have seen a dramatic surge in resource demands. The continuous cycle of tool calls, multi-step reasoning chains, extended context windows, and retry logic has pushed token consumption up significantly in recent years. When these workloads run on a cloud API, every single token incurs a cost. On a locally owned machine, there is no per-token billing — the hardware is purchased once and runs indefinitely.
The new 64GB model fills a specific niche in NVIDIA’s lineup. It retains the same GB10 superchip, CUDA-powered AI software stack, and ConnectX-7 networking capabilities found in the original DGX Spark, but pairs them with 64GB of unified LPDDR5x memory instead of 128GB. NVIDIA positions this configuration as sufficient for running today’s most capable open models in the 30–35 billion parameter range. For developers who need larger single-box capacity, the 128GB variant remains available, and two 64GB units can be clustered together to combine memory and compute power as workloads scale.
## Technical Specifications
At the heart of the system sits the GB10 Grace Blackwell superchip, which marries a Blackwell GPU featuring fifth-generation Tensor Cores with a 20-core Arm architecture CPU. The configuration breaks down as follows:
| Component | Specification |
|—|—|
| Superchip | NVIDIA GB10 Grace Blackwell |
| CPU Architecture | 20-core Arm (10× Cortex-X925 + 10× Cortex-A725) |
| AI Compute | Up to 1 petaFLOP in FP4 with sparsity support |
| Memory | 64GB unified LPDDR5x |
| Memory Bandwidth | 273 GB/s |
| Storage Options | 1TB, 2TB, or 4TB NVMe M.2 with self-encryption |
| Networking | ConnectX-7 NIC at 200GbE, Wi-Fi 7, Bluetooth 5.3 |
| Display Output | 1× HDMI 2.1a |
| Operating System | NVIDIA DGX OS (Ubuntu-based) |
| Physical Dimensions | 150 × 150 × 50.5 mm |
| Weight | 1.2 kg |
| Max Local Model Capacity | Up to 100 billion parameters |
| Launch Date | October 23, 2026 |
| Availability | NVIDIA Marketplace, OEM retail channels |
## Unified Memory Architecture: The Core Innovation
One of the most significant design decisions is the use of unified memory accessible by both the CPU and GPU through NVLink-C2C interconnect. This offers five times the bandwidth of PCIe Gen 5 and eliminates the need to copy model weights between system RAM and dedicated VRAM. For AI agents, this means multiple models, their key-value caches, and associated tool processes can all coexist within a single shared address space.
The practical benefit is substantial. Instead of managing fragmented memory pools across different hardware components, developers can load several models simultaneously and switch between them seamlessly, all while keeping data locality intact.
## Software Ecosystem Readiness
The system ships with a complete software stack out of the box. DGX OS includes PyTorch, Jupyter notebooks, and Ollama for model serving. NVIDIA NemoClaw can be installed with a single command and adds privacy controls and security layers to agent frameworks built on OpenClaw. The NVIDIA OpenShell, part of the broader NVIDIA Agent Toolkit, provides policy-based guardrails that govern agent behavior. Additionally, NVIDIA’s own Nemotron model family is optimized specifically for this hardware platform.
Perhaps most importantly, the entire system is designed to operate on a standard wall outlet. There is no need for dedicated server rooms, specialized cooling infrastructure, or industrial-grade power systems — making it practical to run as an always-on workstation in an office or home lab environment.
## Open Models That Fit Within 64GB
Having adequate hardware is only half the equation. The other half is that open models in the 30 billion parameter class have matured significantly in terms of quality and agent-readiness. Three notable models that run well on the 64GB configuration include:
**Muse Glimmer** by Meta — a 29.6 billion parameter dense model that handles both text and image inputs under an Apache 2.0 license. Distilled from the Muse Spark family that powers Meta’s AI assistant, Glimmer is purpose-built for local agent applications. It delivers strong performance on coding benchmarks (51.2 on SWE-Bench Pro) and multi-step tool-calling scenarios (75.5 on MCP Atlas), with support for context windows up to 131,000 tokens. In its quantized form, the model requires approximately 17GB of memory, leaving ample headroom for context and runtime overhead.
**Nemotron 3.5 Lightning** by NVIDIA — a 30 billion parameter mixture-of-experts model with a 30B active parameter configuration. This model is designed as a fast executor for long-running agent workflows, and its checkpoint fits within the NVFP4 precision format supported by the GB10 superchip.
**Qwen 3.8-27B** by Alibaba — a 27 billion parameter dense model that, when quantized to 4-bit precision, requires roughly 13.5GB for weights alone. This makes it an excellent choice for general-purpose agent tasks and code generation.
For larger models that exceed the single-node capacity, NVIDIA has demonstrated running DeepSeek V4 Flash across four clustered 64GB DGX Spark units.
## Five Practical Use Cases on a Single System
### 1. An Always-On Personal AI Agent
By installing a framework like Hermes and pairing it with Muse Glimmer or Nemotron 3.5 Lightning, developers can create a persistent agent that monitors their code repositories, runs test suites, tracks new research papers, and drafts pull requests — all while the machine sits idle overnight. Private notes, source code, and communications never leave the local machine.
### 2. Fine-Tuning a Coding Model on Proprietary Codebases
Quantized Low-Rank Adaptation (QLoRA) allows fine-tuning of 70 billion parameter models within the 64GB memory envelope. Developers can train on their own repositories, internal documentation, and historical pull request reviews, producing a coding assistant that understands their specific conventions, architecture patterns, and coding standards. NVIDIA has measured throughput of approximately 18,400 tokens per second for distributed fine-tuning workloads on a single node.
### 3. A Day-One Model Evaluation Platform
When a new open-weight model appears on model repositories, it can be pulled and evaluated immediately using Ollama or vLLM. Developers can run custom benchmark suites, measure time-to-first-token and throughput metrics, and compare results against existing models — all without incurring cloud API charges, even after dozens of re-runs.
### 4. A Multi-Model Agent Team
The unified memory architecture makes it possible to run several models simultaneously from a shared pool. A routing model like Bonsai 2 can direct tasks to a primary reasoning agent like Glimmer, with the combined weight footprint coming to roughly 23GB. The remaining memory accommodates key-value caches, the operating system, and tool execution processes. When a task exceeds local capacity, the agent can route a sanitized query to a larger cloud-based model.
### 5. Edge and Robotics Prototyping
The same CUDA acceleration stack used in the DGX Spark can be applied to fine-tuning vision transformers for specific tasks such as defect detection on manufacturing lines. Models validated locally can then be deployed to edge devices like NVIDIA Jetson modules, creating a seamless development-to-deployment pipeline.
## Scaling Out: Clustering with NVIDIA Sync
Every DGX Spark unit ships with ConnectX-7 networking capable of 200GbE connections. Clustering is a native capability rather than an optional add-on.
| Configuration | Combined Memory | AI Compute (FP4) | Connection Method |
|—|—|—|—|
| 1 unit (64GB) | 64GB | Up to 1 PFLOP | — |
| 2 units (64GB each) | 128GB | Up to 2 PFLOPS | Direct QSFP cable, no network switch required |
| 2 units (128GB each) | 256GB | Up to 2 PFLOPS | Direct QSFP cable, no network switch required |
| 3 units (128GB each) | 384GB | Up to 3 PFLOPS | QSFP ring topology |
| 4 units (128GB each) | 512GB | Up to 4 PFLOPS | 200GbE switch with QSFP56-DD and RoCE v2 |
NVIDIA reports that two clustered 64GB units deliver up to 1.7 times the performance of a single 128GB DGX Spark, owing to the doubling of both AI compute and memory bandwidth — the 64GB configuration achieves up to 546 GB/s combined bandwidth across two nodes.
The NVIDIA Sync management application runs on both Windows and macOS, automatically discovering DGX Spark units on the local network and managing SSH connectivity. The Cluster Assistant handles ConnectX-7 networking configuration for up to four systems. Additionally, nodes can be linked across physical locations through a Tailscale mesh network without requiring cloud infrastructure in the data path.
The performance scaling characteristics are worth noting: clustering approximately halves time-to-first-token for every doubling of nodes, and fine-tuning throughput scales nearly linearly. Decode performance improvements are more modest, reaching roughly 1.4 times at four nodes. For agent workloads that process long input sequences, the reduction in time-to-first-token provides the most meaningful performance gain.
Detailed multi-node guides, including vLLM deployment across stacked DGX Spark units, are available through NVIDIA’s developer documentation portal.
## Understanding the System’s Boundaries
The DGX Spark 64GB is a purpose-built workstation, and it is important to understand what it is designed for and what it is not:
– **Not a multi-user chat server.** The 273 GB/s memory bandwidth constrains concurrent decode throughput for many users simultaneously.
– **Optimized for long-input, short-output patterns.** The system excels at reading repositories, log files, document collections, or research papers and producing concise outputs.
– **64GB limits single-node models to approximately 100 billion parameters.** For larger models, clustering two units or selecting the 128GB variant is the recommended path.
## Frequently Asked Questions
**What makes the DGX Spark different from a standard desktop workstation?**
The DGX Spark uses NVIDIA’s GB10 Grace Blackwell superchip, which integrates a Blackwell GPU and Arm CPU on a single module connected by NVLink-C2C. This unified architecture eliminates the traditional bottleneck of moving data between separate CPU and GPU memory pools, providing five times the interconnection bandwidth of standard PCIe Gen 5.
**Can I run this system without an internet connection?**
Yes. The entire software stack, including DGX OS, PyTorch, Ollama, and NVIDIA’s agent tools, is pre-installed and functional offline. Models can be downloaded and stored locally. Clustering over a Tailscale mesh also does not require data to pass through any cloud service.
**How much does the 64GB DGX Spark cost?**
Pricing information for the 64GB configuration is available through NVIDIA’s marketplace page and the various OEM partner retail channels. The 128GB variant continues to be offered alongside the new 64GB model.
**What networking is required for clustering?**
Two DGX Spark units can be connected directly via a QSFP cable with no network switch needed for up to two nodes. For three or four nodes, a 200GbE switch supporting QSFP56-DD and RoCE v2 is required. The NVIDIA Sync application simplifies the entire setup process.
**Is the 64GB model suitable for training large language models from scratch?**
The system is better suited for fine-tuning existing models using techniques like QLoRA rather than training foundation models from scratch. However, it provides substantial compute power for distributed training when multiple units are clustered together.
**Which operating systems are supported?**
The system runs NVIDIA DGX OS, which is based on Ubuntu. The NVIDIA Sync management application is available for both Windows and macOS hosts.
**How does the 64GB model compare to the 128GB model for single-node workloads?**
The 64GB model uses the same GB10 superchip and delivers identical AI compute performance (1 PFLOP FP4). The only difference is memory capacity — 64GB versus 128GB — which affects how large a single model can be and how much room remains for context windows and runtime overhead.
## Final Thoughts
The introduction of the 64GB DGX Spark represents a significant shift in how developers can approach AI agent development. By offering a compact, desk-friendly workstation that runs the most capable open models locally, NVIDIA is enabling a workflow where agents operate continuously without recurring cloud costs, sensitive data never leaves the premises, and scaling from a single node to a multi-unit cluster is a matter of plugging in additional systems.
As token consumption continues to climb and open models grow more capable, having a local hardware foundation that eliminates API dependency becomes increasingly valuable — both from a cost perspective and from a perspective of control, privacy, and reliability.
The 64GB DGX Spark is available for purchase starting October 23, 2026, through NVIDIA’s official marketplace and all major OEM partners.
Thank you for reading



