NVIDIA and Microsoft used a Windows and Surface event on October 7 to outline a new generation of PCs intended to run substantial AI models and persistent agents locally. The centerpiece is NVIDIA RTX Spark, a combined Grace CPU and Blackwell RTX GPU platform for laptops and compact desktops. Laptop preorders opened at the event, with availability scheduled for October 16, while smaller desktop systems are expected in November. The companies also previewed a much larger DGX Station for Windows.

RTX Spark combines up to a 20-core Grace CPU with a Blackwell GPU offering as many as 6,144 cores. NVIDIA specifies 600 GB/s of connectivity between the processor components, up to 128 GB of unified memory and a peak of one petaflop of FP4 AI compute. Those are maximum configuration and low-precision performance figures, not guarantees for a particular model or workload. NVIDIA says the system can run large models locally without sending data to a cloud inference service, a proposition that will appeal to developers handling private code and documents.

Microsoft introduced Surface Laptop Ultra as one of the RTX Spark systems, alongside designs from Acer, ASUS, Dell, HP, Lenovo, MSI and Gigabyte. The Verge reports that Surface Laptop Ultra starts at $2,599 in the United States, with higher-end configurations reaching about $5,900. The base configuration does not include the headline 128 GB memory capacity; buyers need to distinguish the platform's maximum specification from the hardware actually offered at each price. PCWorld's hands-on coverage also shows a broad and expensive first wave of devices.

The software announcement is equally important for agent deployment. Microsoft said Microsoft Execution Containers, or MXC, are generally available as an operating-system mechanism for running agents persistently under Windows controls. NVIDIA presents the combination of local acceleration, CUDA tooling and operating-system isolation as a way to move agent workloads from remote servers onto personal or enterprise machines. The companies did not demonstrate that local execution is always cheaper or safer; those outcomes depend on the model, energy use, access controls and workload.

For developers requiring more capacity, NVIDIA previewed DGX Station for Windows. It uses a GB300 Grace Blackwell Ultra Desktop Superchip with 748 GB of coherent memory and up to 20 petaflops of FP4 compute, according to NVIDIA. The company says this class of system could accommodate models approaching a trillion parameters locally. The preview is not a general-availability announcement, and a system's ability to load a model is different from meeting useful inference throughput and latency targets.

This is a notable attempt to make local AI a first-class Windows computing category, rather than an occasional accelerator feature. The unresolved questions are pricing beyond the initial premium tier, real sustained performance under laptop thermal limits, compatibility with existing agent stacks, and whether organizations can secure agents that have continuous access to local applications and files. The hardware specifications and roadmap come from NVIDIA and Microsoft; independent reporting confirms the product announcements but not every performance projection.

References