Structure, Symmetry, and Scaling in Neural Network Weight Space
by Daniel Herbst
Modern neural networks are parameterized by high-dimensional weight vectors, but these representations are far from unique, i.e., many different points in weight space can describe the same function. Such weight-space symmetries are relevant across machine learning. For example, they shape the geometry of loss landscapes, can induce conserved quantities in optimization dynamics, and they are obstructions that have to be accounted for when comparing or merging independently trained models. Conversely, moving along their orbits can create opportunities for improved compression (e.g., quantization or pruning), and they can be leveraged as data symmetries for weight space learning.
My research studies this structure in weight space. One focus is on moving beyond purely architectural symmetries, but considering more local and data-dependent notions of equivalence, i.e., transformations that may not preserve a network’s function everywhere, but preserve its behavior on the inputs and representations that are relevant in practice. A complementary question is how such notions of equivalence extend across dimensions. This shifts the problem from comparing points within a single weight space to comparing models of different widths or depths that inhabit different weight spaces. Understanding such correspondences may clarify which properties of neural networks persist under scaling and which depend on a particular model size or parameterization. Ultimately, I seek mathematical descriptions of neural networks that capture what remains invariant across parameterizations, architectures, and scales.
Vincent Bürgin*, Daniel Herbst*, Ya-Wei Eileen Lin, and Stefanie Jegelka (2026): Beyond Structural Symmetries: Linear Mode Connectivity via Neuron Identifiability. International Conference on Machine Learning (ICML), 2026
Daniel Herbst and Stefanie Jegelka (2025): Higher-Order Graphon Neural Networks: Approximation and Cut Distance. International Conference on Learning Representations (ICLR), 2025
