Microsoft says its newest Surface can run AI models exceeding 120 billion parameters locally, depending on configuration.
Microsoft has opened preorders for the Surface Laptop Ultra, with the first systems shipping Oct. 16. The laptop starts at $2,599 and is aimed at developers, creators and users who want to run demanding AI workloads locally rather than sending every task to the cloud.
At the heart of the machine is Nvidia’s RTX Spark platform, combining a Grace CPU with a Blackwell GPU. Higher configurations offer up to 20 CPU cores, 6,144 GPU cores and 128GB of unified memory. Microsoft claims up to 1 petaflop of AI performance and says the system can run models exceeding 120 billion parameters locally.
For channel partners, the opportunity is helping customers determine which workloads justify the hardware investment, how much memory they need, and whether their applications work on Windows on Arm.
Surface Laptop Ultra: Key specs
| Feature | Details |
| Starting price | Starting at $2,599 MSRP; pricing varies by configuration |
| Display | 15-inch PixelSense Ultra touchscreen (3270 x 2180, 3:2, 120Hz, 2,000 nits peak HDR) |
| Processor | Nvidia RTX Spark N1X, up to 20-core CPU |
| GPU | Up to 6,144 Blackwell GPU cores |
| AI performance | Up to 1 petaflop theoretical FP4 AI compute using sparsity; Nvidia NPU |
| Storage | Removable 512GB–2TB |
| Unified memory | 24GB, 32GB, 48GB, 64GB, or 128GB unified LPDDR5x RAM |
| Weight | About 4.41 pounds |
| Thickness | 17.99mm without feet; 19.16mm with feet |
| Charging | Magnetic Connect on primary USB-C (140W fast-charging compatible) |
| Wireless | Wi-Fi 7, Bluetooth 5.4 |
| Operating system | Windows edition varies by configuration; Surface Laptop Ultra for Business ships with Windows 11 Pro |
| Battery | 92 Wh battery (Up to 15 hours video playback, 12 hours web usage) |
| Ports | 3 USB-C, HDMI, USB-A, SD, 3.5mm |
| Warranty | One-year limited hardware warranty |
| In the box | Laptop and USB-C Magnetic Connect cable; 140W adapter inclusion varies by market and configuration |
Industrial design, display, and breakaway USB-C
Beyond raw silicon, Microsoft introduced Magnetic Connect, which it describes as the first built-in magnetic USB-C charging port on a laptop. Rather than relying on a proprietary charging standard, Microsoft built a magnetic breakaway mechanism directly into an open-standard USB-C port.
The bundled 140W power cable snaps free if pulled accidentally, but the port itself remains fully compatible with regular third-party USB-C cables, external peripherals, and 40 Gbps data transfers.
The 15-inch PixelSense Ultra screen features a 3:2 aspect ratio at 3270 x 2180 resolution (262 PPI) with dynamic 120Hz refresh rates. Microsoft claims peak HDR brightness reaches 2,000 nits, which Microsoft says is up to 25% higher than the MacBook Pro M5 Pro’s published peak HDR specification. However, Microsoft noted that the screen is optimized exclusively for 10-point touch interaction; it does not support Surface Pen or Surface Slim Pen input.
Connectivity includes two additional USB-C ports supporting DisplayPort 2.1, full-size HDMI 2.1b, USB-A 3.1, a full-size SD card reader, and a 3.5 mm audio jack. The SSD is removable and can be replaced or upgraded following Microsoft’s service guidance.
Alongside the laptop, Microsoft opened preorders for the Surface RTX Spark Dev Box at $5,999, a compact desktop developer workstation housing the same 128 GB RTX Spark silicon, shipping to US customers in November.
Software architecture and agent sandboxing
The machine arrives during an architectural shift for Windows 11, moving from static chat assistants to autonomous background agents. To govern local automated workflows, Microsoft announced the general availability of Microsoft Execution Containers (MXC).
Agents running with a user’s permissions can access resources beyond those needed for a task. MXC lets developers define file and network access policies enforced through a selected containment backend. Hardware-enforced isolation is provided by its experimental microVM option.
The policies operate independently outside the model’s environment so software agents cannot escalate their own privileges. Microsoft confirmed that GitHub Copilot, OpenAI Codex, LM Studio, and Replit are adopting the system, with Anthropic’s Claude Code and Perplexity integrating support soon.
Microsoft plans to bring HydraFusion’s local-and-cloud model routing to GitHub Copilot in experimental preview later in October. Separately, it announced llama.cpp support in Windows ML.
Contextual attributions and competitive claims
During the keynote, leadership leaned heavily into comparisons against Apple’s silicon hardware. Pavan Davuluri, Microsoft’s executive vice president of Windows and devices, framed the machine as a bridge to autonomous workflows, writing that “we’re building Windows as the home for hybrid intelligence: a platform where agents can run locally when it makes sense.”
Nvidia CEO Jensen Huang joined Nadella to discuss their companies’ work on Windows PCs designed to run AI agents.
Huang later added that “MXC is going to revolutionize how agents are built and deployed,” while Microsoft CEO Satya Nadella told attendees, “We needed to make the desktop the most secure place for agents to execute.”
To lure corporate and creative users away from Apple, Microsoft is offering up to $1,000 cash back with an eligible MacBook Pro trade-in and qualifying Surface purchase through Nov. 23, available through Microsoft Store online in the US and Canada. Microsoft reports up to 6.2 times faster AI video generation and 2.1 times faster time to first token across tested preproduction RTX Spark PCs versus a 16-inch MacBook Pro with M5 Pro and 64GB of memory. Its published footnotes identify the models, workloads, and test settings, but these vendor-reported results do not establish performance across every application or configuration.
The true cost of local independence
Running supported workloads locally can reduce cloud inference spending and limit the data sent to external services. However, the Microsoft and Nvidia partnership does not eliminate ongoing costs: customers still need to budget for software, electricity, maintenance, security, and workloads that continue to use cloud services.
The base model’s 24GB of memory should not be confused with the platform’s maximum advertised model capacity. Memory requirements depend on the model, quantization, context window, and other running applications. Partners should weigh higher-memory hardware costs against expected cloud usage before recommending a configuration.
What this means for Windows buyers and channel partners
For Windows buyers, the decision is simple: if local AI is central to your workflow, the Surface Laptop Ultra offers capabilities that ordinary premium laptops do not. If most of your work is Office, browsing, meetings and conventional productivity, its AI hardware is difficult to justify at this price.
Channel partners have an opportunity to sell the machine around workloads rather than specifications. Customers will need help determining how much unified memory they actually require, which AI models can run locally and whether their software, drivers and plugins work properly on its Arm-based platform.
Before recommending a preorder, partners should test a representative customer workload, confirm application and driver compatibility, and compare deployment and support costs with continued cloud use. That assessment will show whether the hardware premium delivers value for the customer.
Read more: Explore how local AI could help MSPs build faster pilots and hybrid AI services as customers weigh cost, privacy, and performance.



