Nvidia launches DGX Spark 64 GB for $4,999 to run AI agents and models locally
Nvidia announced on October 2 2026 a version of its DGX Spark platform with 64 GB of unified memory, available from October 23 at a starting price of $4,999 through various hardware manufacturers.

Announcement and availability
On October 2 2026, Nvidia officially announced the launch of a new configuration within its DGX Spark family, named DGX Spark 64 GB. This model incorporates a unified memory pool of 64 GB, merging system RAM and GPU VRAM into a single addressable space. The product will become available for purchase starting October 23 2026, with a base price set at $4,999, positioning it as a cost‑effective entry point for on‑premises AI workloads. Distribution will be handled through a network of established channel partners, including Acer, ASUS, Dell, Gigabyte, HP and MSI, giving customers the flexibility to select from a range of well‑known reference brands. Nvidia markets this configuration as the most affordable option for running AI agents and performing fine‑tuning of models without relying exclusively on cloud services, emphasizing lower total cost of ownership and data proximity.
From an architectural standpoint, the DGX Spark 64 GB retains the same GB10 Grace Blackwell super‑chip found in the 128 GB variant, ensuring comparable raw compute capability. The system runs on the dedicated DGX OS operating system, which provides an optimized software layer for AI workloads. Network connectivity is delivered via the ConnectX‑7 interface, offering high‑speed links, while the CUDA development stack remains fully integrated, allowing developers to exploit GPU parallelism efficiently. The unified memory architecture combines RAM and VRAM into a single address space, simplifying data movement between the central processing unit (CPU) and the graphics processing unit (GPU) and reducing overhead associated with data copies. Grace Blackwell’s presence guarantees compute performance approaching a petaflop, and ConnectX‑7 enables multi‑node configurations with very high‑bandwidth interconnects.
Model capacity and performance
Nvidia claims that a single DGX Spark 64 GB unit can run models up to 100 000 million parameters, opening the door to fairly large neural networks while staying within the unified memory limits. When two units are linked together via a 200 GbE network, the combined memory rises to 128 GB, effectively doubling the maximum model size to 200 000 million parameters. Internal testing with the Qwen 3.8 27B model demonstrated that a two‑unit configuration delivered roughly a 1.7× performance uplift compared with a single unit, indicating efficient scalability for compute‑intensive workloads. These results suggest that the system can be scaled linearly, at least for the tested scenarios, while maintaining acceptable energy efficiency and latency characteristics.
The capacity and performance figures originate directly from Nvidia. The MarkTechPost article that covered the announcement notes that the piece is sponsored and that the numbers should be regarded as supplier statements rather than third‑party verified results. This disclaimer is important for readers seeking an independent assessment of the hardware, as it highlights that the metrics have not yet undergone external validation or public benchmarking.
Intended use cases
According to Nvidia, the DGX Spark 64 GB is designed to run continuous agents, perform real‑time inference, conduct model fine‑tuning, and support data‑science workflows without requiring a constant cloud connection. Reduced latency, stemming from the physical proximity of data and compute, together with enhanced local data privacy, are presented as key arguments for sectors such as healthcare, finance and manufacturing, where protecting sensitive information and delivering rapid responses are critical. By processing data on‑premises, organizations can avoid network bottlenecks, keep critical information under direct control, and meet regulatory requirements that mandate data residency.
Nevertheless, the 64 GB unified memory capacity imposes limits on the context length that models can handle and on the number of concurrent requests the system can serve. For applications that demand extensive context windows or high concurrency, Nvidia recommends scaling out by interconnecting multiple DGX Spark units. This approach naturally raises the overall cost of ownership and adds architectural complexity, as it requires high‑bandwidth networking, more sophisticated resource management, and longer‑term capacity planning.
Comparison with the competition
In the market for on‑premises AI hardware, other vendors offer solutions based on consumer‑grade GPUs or specialized accelerators, but these alternatives often come at higher price points or with less integrated software stacks. The DGX Spark combines top‑tier hardware with an optimized software stack (DGX OS and CUDA), potentially reducing integration time for development teams. With a starting price of $4,999, the system sits below most reference configurations that exceed $8,000, although direct performance comparisons depend heavily on the specific workload and the client’s performance requirements.
For organizations operating in English‑speaking markets, the arrival of the DGX Spark 64 GB opens the possibility of running advanced AI projects locally without relying solely on cloud services. Companies can keep sensitive data within their own data centers, lower data transfer costs, and take advantage of an architecture that allows gradual scaling by adding additional nodes. In practice, this translates into greater technological autonomy, a reduced carbon footprint associated with data traffic, and an accelerated adoption of AI agents in critical processes where speed and confidentiality are paramount.
Sources
- NVIDIA DGX Spark 64GB Gives Developers More Ways to Build and Scale Local AINVIDIA · October 2, 2026
- NVIDIA Announces DGX Spark 64GBMarkTechPost · October 2, 2026



