DGX Spark · documentation

Sources

SOURCES.md Last pushed 19 August 2026

DFlash 2

Published figures cited for comparison (Inco AI, H200, SGLang, block size 8, temperature 1.0 with xhigh reasoning): acceptance length 4.80 mean for Qwen3.8-27B against MTP 4.28 and a community DSpark drafter 3.62; throughput 2.7-3.4x autoregressive at concurrency 1.

Models

Engines

Evaluation data

Prior art referenced by DFlash 2

  • Canon Layers; Dynamic Short Convolutions; Convolution for Large Language Models — cited in the announcement as the basis for the two-tap dynamic depthwise convolution.
  • Modal, Speculation Is All You Need — cited on speculative decoding for low-latency serving.

Every number here was measured. Open an issue if one looks wrong.

All documentation