Pay Attention to Attention Distribution: A New Local Lipschitz Bound for Transformers
Abstract
We introduce a novel upper bound on the local Lipschitz con-stant of the dot-product self-attention block showing its dependence onthe attention map distributions. The proposed bound is not only tighterthan the prior art, but for the first time, reveals how the distributionof attention probabilities shapes the local Lipschitz constant of the self-attention block. The theoretical basis of the proposed upper bound lies inthe refined closed-form upper bounds on singular values of the Jacobianof softmax function. Leveraging these theoretical insights, we introduceJaSMin (Jacobian Softmax norm Minimization), a lightweight regular-izer that directly controls the local Lipschitz constant of each block and,consequently, the entire model. Additionally, we discuss how the natureof the attention map distribution contributes to the gradient dynamicsand, consequently, transformer training stability.