NVIDIA’s 64GB DGX Spark brings larger models and agents to local systems

NVIDIA says its 64GB DGX Spark can run 100B-parameter models and local agents, while two systems can be synchronized into a 128GB setup.

NVIDIA announced a 64GB unified-memory configuration of DGX Spark on October 2, 2026, moving the compact AI development system further toward local model and agent workloads. NVIDIA says the new configuration can run models at the 100B-parameter scale and support coding, research, and other agent workflows without depending on the cloud. Partner systems are scheduled to start at $4,999 on October 23.

The point of DGX Spark is not simply putting a GPU in a smaller enclosure. It combines GB10, 64GB of unified memory, DGX OS, and NVIDIA's software stack so models, tools, and agents can run on one local system. Local processing can reduce the need to send sensitive code, internal documents, or offline data across a network, but privacy still depends on networking, logs, user permissions, and installed tools. A local label is not a complete security control.

When one system is not enough, two DGX Spark units can use NVIDIA Sync Cluster Assistant to form a 128GB local system, which NVIDIA says can handle models up to roughly 200B parameters. NVIDIA also reports up to 1.7 times the performance of one system when two synchronized units run Qwen 3.8 27B in its testing. That is a vendor hardware and software result; real performance will depend on quantization, networking, batch size, parallelism, and workload shape.

The software stack includes or supports NVIDIA Agent Toolkit, CUDA-X AI, Nemotron, Ollama, vLLM, and PyTorch. Sync helps configure networking, software versions, and nodes, while a future launcher is expected to add Qwen 3.8 and OpenCode. That integration matters for local agents because the job is not just downloading a model. Teams also need to manage inference servers, tool permissions, data directories, task state, and coordination across machines.

NVIDIA lists use cases such as running a coding or research agent around the clock, handling application inference on a local PC, and scaling from one system to two as workload grows. Local agents can give teams more direct control over data and recurring cloud costs, but the tradeoff is operational responsibility for hardware supply, updates, backups, monitoring, access control, and failure recovery.

The 64GB DGX Spark announcement makes local execution a clearer option for AI agents. Not every task needs a large cloud service, and not every local workload will achieve the same model scale or speed. Before buying, teams should measure context length, concurrency, quantization, latency, and power use, then decide which data belongs locally. For production, the local environment still needs to be operated and audited as a service; it is not a secure black box simply because it sits nearby.

MODULE.002 //

More insights

Ideas on websites, AI automation, digital marketing, AI news, and VMTS updates.