Nvidia pushes CXL to cut inference costs

Nvidia and partners are promoting CXL and offload engines such as CMX to move key-value caching to SSDs and DPUs, aiming to lower AI inference economics.

Mateo Fernandez ·

Nvidia pushes CXL to cut inference costs

Nvidia is advancing CXL-based architectures and offload engines to reduce the cost of AI inference, a briefing dated July 17, 2026 said. The note describes using CMX and CXL to keep XPUs fed by shifting key-value (KV) caches from memory into SSDs and DPUs; the company announced the approach intends to cut per-query compute and memory pressure. Market reaction pending.

CMX and CXL offload engines

The proposal routes KV cache traffic over CXL links so SSDs and smart NICs with DPUs handle large, cold caches while XPUs keep hot working sets local. Officials said this separates capacity storage from compute bandwidth, letting vendors scale cheaper flash capacity without replicating DRAM across every accelerator.

Engineers argue the model reduces the need for oversized HBM and large memory-coherent pools on every accelerator, shifting capital and operational costs toward higher-capacity SSDs and programmable DPUs. Data showed system designers are testing latency and throughput trade-offs to ensure inference SLAs are met.

Adoption could reshape server designs and procurement: hyperscalers may prioritize CXL-capable infrastructure and vendors that supply integrated SSD+DPU offload stacks. Officials said software and firmware maturity will determine how quickly customers accept the latency trade-offs.

Expect follow-up technical disclosures and vendor pilots by July 24, 2026 as companies outline performance and cost benchmarks.

More stories