Devices / NVIDIA Grace Blackwell · GB10 (product GX10)

DGX Spark · NVIDIA GB10

A measured AArch64 build and execution host whose accelerator is not yet admitted by any Alpha device lens.

AArch64 host · Blackwell compute capability 12.1

Qualification

Host yes · accelerator no

Emitted AArch64 programs ran here and the host facts were read from the machine (2026-09-18). The GB10 accelerator is rejected by the SM86 capability lens on purpose: no Alpha owner has demonstrated QMD, pushbuffer or GPFIFO submission on it.

Driver facts · live

Typed profile

Read at build time from Platform.DGXSpark.Profile at 408ac92b. Systems bind to these values by name; none branches on a card name.

FactValueAlpha definition
Profile identitydgx-spark-bringup-v1dgxSparkProfileIdentity
Performance cores10 × Cortex-X925dgxSparkPerformanceCoreCount
Efficiency cores10 × Cortex-A725dgxSparkEfficiencyCoreCount
Unified memory130,596,048,896 B (121.6 GiB)dgxSparkTotalMemory
Page size4096 BdgxSparkPageSize
Level 2 cache26,214,400 B (25 MiB)dgxSparkLevel2Cache
Level 3 cache25,165,824 B (24 MiB)dgxSparkLevel3Cache
Accelerator compute capability12.1dgxSparkDeviceComputeMajor / Minor

Profile coverage

What has been measured, and what has not

Three tiers make a fully profiled device. A tier marked not measured lists exactly the fields it will fill so the gap is explicit.

  1. 1

    Driver facts

    Live

    What the host reports: core parts and counts, unified memory, page size, caches, and the accelerator's compute capability as seen through the SM86 lens.

    • ARM implementer and core parts
    • core counts
    • unified memory
    • page size
    • L2 / L3 cache
    • GB10 compute capability
  2. 2

    Micro-benchmarks

    Not measured yet

    Measured instruction and memory behaviour, so the compiler can reason about cost from numbers read on this card rather than assumptions.

    • issue rate and latency per typed instruction (FFMA, IMAD, MUFU, LDG, STG, LDSM, HMMA…)
    • shared-memory bandwidth and bank behaviour
    • global-memory bandwidth by access pattern and width
    • tensor-core throughput at 16×8×16
    • barrier and warp-shuffle cost
    • submission, doorbell and semaphore round-trip cost
    • sustained clocks under the training workload
  3. 3

    Model-shaped calibration

    Not measured yet

    The cost of the exact kernels and launch schedules Alpha emits for each system on this device, so realization choices can be made from measurements.

    • per-kernel time for every launch in Bob's and Coppelius's schedules
    • occupancy and register / shared-memory pressure per kernel
    • end-to-end step time and its split between host protocol, launches and waits
    • memory footprint against the profile's video memory

Sources

Pairings and evidence