Devices / NVIDIA Grace Blackwell · GB10 (product GX10)
DGX Spark · NVIDIA GB10
A measured AArch64 build and execution host whose accelerator is not yet admitted by any Alpha device lens.
AArch64 host · Blackwell compute capability 12.1
Qualification
Host yes · accelerator no
Emitted AArch64 programs ran here and the host facts were read from the machine (2026-09-18). The GB10 accelerator is rejected by the SM86 capability lens on purpose: no Alpha owner has demonstrated QMD, pushbuffer or GPFIFO submission on it.
Driver facts · live
Typed profile
Read at build time from Platform.DGXSpark.Profile at 408ac92b. Systems bind to these values by name; none branches on a card name.
| Fact | Value | Alpha definition |
|---|---|---|
| Profile identity | dgx-spark-bringup-v1 | dgxSparkProfileIdentity |
| Performance cores | 10 × Cortex-X925 | dgxSparkPerformanceCoreCount |
| Efficiency cores | 10 × Cortex-A725 | dgxSparkEfficiencyCoreCount |
| Unified memory | 130,596,048,896 B (121.6 GiB) | dgxSparkTotalMemory |
| Page size | 4096 B | dgxSparkPageSize |
| Level 2 cache | 26,214,400 B (25 MiB) | dgxSparkLevel2Cache |
| Level 3 cache | 25,165,824 B (24 MiB) | dgxSparkLevel3Cache |
| Accelerator compute capability | 12.1 | dgxSparkDeviceComputeMajor / Minor |
Profile coverage
What has been measured, and what has not
Three tiers make a fully profiled device. A tier marked not measured lists exactly the fields it will fill so the gap is explicit.
- 1
Driver facts
LiveWhat the host reports: core parts and counts, unified memory, page size, caches, and the accelerator's compute capability as seen through the SM86 lens.
- ARM implementer and core parts
- core counts
- unified memory
- page size
- L2 / L3 cache
- GB10 compute capability
- 2
Micro-benchmarks
Not measured yetMeasured instruction and memory behaviour, so the compiler can reason about cost from numbers read on this card rather than assumptions.
- issue rate and latency per typed instruction (FFMA, IMAD, MUFU, LDG, STG, LDSM, HMMA…)
- shared-memory bandwidth and bank behaviour
- global-memory bandwidth by access pattern and width
- tensor-core throughput at 16×8×16
- barrier and warp-shuffle cost
- submission, doorbell and semaphore round-trip cost
- sustained clocks under the training workload
- 3
Model-shaped calibration
Not measured yetThe cost of the exact kernels and launch schedules Alpha emits for each system on this device, so realization choices can be made from measurements.
- per-kernel time for every launch in Bob's and Coppelius's schedules
- occupancy and register / shared-memory pressure per kernel
- end-to-end step time and its split between host protocol, launches and waits
- memory footprint against the profile's video memory
Sources