Φeat: Physically-Grounded Material Feature Representation
Abstract
While foundation models have emerged as general-purposevisual backbones, their representations are primarily optimized for se-mantics and lack explicit modeling of physical factors, such as reflectance,hindering their efficacy in tasks requiring explicit material reasoning. Weintroduce Φeat, a novel material-grounded visual backbone that encour-ages a representation sensitive to material identity, including reflectanceand mesostructure. Instead of relying on generic data augmentations,we pretrain our model by contrasting observations of the same materialunder controlled variations in lighting and geometry. This encouragesinvariance to extrinsic factors while preserving sensitivity to intrinsicmaterial properties. We show that the resulting representation providesstrong priors for material-centric tasks, including feature-based mate-rial selection and classification. Our results demonstrate that physicallyinspired weak supervision is an effective strategy for learning represen-tations tailored to material perception.