All benchmarks

GPU and enclave model

tasqnetwork.io/benchmark/gpu-model

Effective mode A overhead by job length

steady-state 2%steady-state 5%steady-state 7%
10%100%0.10.521060session length (s, log scale)effective overhead (log scale)

Session setup takes about 120 ms in this model, of which 0.44 ms is cryptography. Short jobs should share a session. Past about a minute the overhead is the steady-state figure. Steady-state range from Zhu et al., arXiv:2409.03992.

Source: M1_modeA_overhead.csv, model_gpu.py

When is mode A cheaper than redundancy?

r = 2r = 3r = 5
0x1x2x3x4x5x6x7x0.000.050.100.150.200.250.30audit rate πbreak-even attested premium

Mode A is cheaper than mode R while attested GPU time costs less than this multiple of commodity GPU time. For r = 3 the break-even is about 3x, so confidentiality does not have to cost more than redundancy.

Source: M3_breakeven_premium.csv, M2_mode_costs.csv

Floating-point reordering

ReorderingBit-identical rowsMedian max rel. diffTop-1 agreementTop-5 set agreementSamples
reversed blocks0%3.75e-7100%100%400
interleaved0%3.75e-7100%100%400

Changing only the order of float32 accumulation in a two-layer network breaks bit identity on every row while the predicted tokens stay the same. This is why mode R compares outputs within a tolerance instead of byte for byte. It runs on a CPU and stands in for, but does not measure, GPU kernel nondeterminism.

Source: D1_nondeterminism_cpu.csv