Codes and Expansions (CodEx) Seminar


David Heffren (Johns Hopkins)
Optimizing dimension to learn any-dimensional models

Many data modalities come naturally in varying sizes, for example point clouds, graphs, and text sequences. Models designed for such inputs, including graph neural networks (GNNs) and deepsets, can take inputs of different sizes or dimensions in a way that is transparent to the user, and exhibit size generalization (or transferability) under suitable assumptions. These are often called any-dimensional machine learning models. However, the role of input dimension in shaping learning dynamics and optimization is still poorly understood. In this work, we study how the dimension of the training data influences the optimization through the lens of biased gradient descent. We propose a general framework for analyzing this effect and introduce an algorithm that adaptively selects the training data dimension during the training dynamics for generic any-dimensional models. We prove that the proposed algorithm is approximately optimal for amount of descent per compute. We illustrate our approach on point cloud and graph learning tasks, demonstrating the effectiveness of our algorithm.