Fuzzy PyTorch: numerical variability in deep learning
Fuzzy PyTorch combines an instrumented PyTorch build with Verificarlo and the PRISM backend to evaluate floating-point variability in deep-learning computations. I created the PRISM library, contributed to Fuzzy PyTorch, and coauthored the TMLR paper with Inés Gonzalez Pepe, Hiba Akhaddar, and Tristan Glatard.
Which uncertainty is being evaluated?
For a fixed model \(f_\theta\) and input \(x\), write an instrumented inference result as \(Y_r=f_\theta(x;\omega_r)\). Here \(\omega_r\) changes arithmetic realizations; it does not denote a new set of learned weights. Hold the checkpoint, preprocessing, dropout configuration, and application-level random seeds fixed when isolating arithmetic sensitivity.
The versioned documentation describes stochastic rounding (SR) and up-down rounding (UD) configurations. Distance-weighted SR and equal-probability UD do not have the same error distribution. The rounding guide derives a small example of this distinction; implementation-specific coverage and rounding behavior must still be checked in PRISM.
Execution and analysis workflow
The reviewed repository targets Linux and provides build recipes for instrumented containers. Select the arithmetic mode, build the documented PyTorch environment, and run an existing inference or training script inside it. Additional packages must preserve the instrumented PyTorch and BLAS/LAPACK builds; the upstream instructions discuss installation with dependencies disabled for this purpose.
Capture predictions before taking an argmax or threshold. A discrete classification can remain constant while the underlying scores vary. For segmentation, examine both spatial disagreement and the downstream regional measurements. For training, distinguish arithmetic perturbations from changes in initialization, minibatch order, and nondeterministic kernels.
For repeated scalar predictions, estimate
\[\bar y=\frac{1}{R}\sum_{r=1}^{R}y_r,\qquad s_y^2=\frac{1}{R-1}\sum_{r=1}^{R}(y_r-\bar y)^2.\]The standard deviation \(s_y\) characterizes individual realizations. The standard error \(s_y/\sqrt R\) characterizes the estimated mean under independent sampling. Reporting the latter as numerical variability would understate the dispersion of a single execution.
Minimal analysis example
Download and run the rounding experiment:
python3 stochastic_rounding.py
This standard-library example isolates rounding probabilities from a neural network’s other components. It uses exact rational arithmetic between fixed-grid rounding steps, so its expected mean and variance can be derived independently of the implementation. It is an introductory diagnostic, not an instrumented PyTorch benchmark.
For actual models, use the pinned container recipes and experimental workloads. Those workloads include MNIST, FastSurfer, and WavLM. Reproducing their timings requires the corresponding build, inputs, and execution environment.
Findings and limits
The TMLR paper reports runtime reductions of 5–60 times relative to Verrou across its evaluated configurations and models spanning 1–341 million parameters. These are comparisons for the paper’s tested workloads, not a speedup guarantee for another architecture. See the research summary for the study’s method and interpretation.
The instrumented scope determines which sources of arithmetic variability are observable. A CPU-oriented build does not establish instrumentation of every GPU operation. Model accuracy, arithmetic dispersion, and uncertainty about the data-generating process are different quantities and should be evaluated separately.