Art Beyond Semantics: Sheaf-Informed Contrastive Learning for Multi-Relational Representations
Abstract
Understanding a painting is never a single act. Art histori-ans may analyze the same work through concepts of style, iconography,or historical context, dimensions that are not interchangeable, and eachcarries distinct semantic relationships between the visual and the tex-tual. Vision-Language Models (VLMs) like CLIP, which learn a singleshared embedding space, collapse this richness into a single homogeneousalignment, thereby losing the multi-relational structure that defines art-historical reasoning. We introduce CANVAS (Contrastive Art-awareNetwork for Vision-Language Alignment with Sheaves), a frameworkfor learning relation-aware multimodal representations inspired by sheaftheory. Each artwork is projected into multiple embeddings conditionedon the type of relation (i.e., the context), and a novel contrastive lossencodes contextual information during training, with no dependency onexternal data at inference. We evaluate on three newly introduced bench-marks of artworks for multi-relational art understanding: WikiArt+,derived from WikiArt and Wikipedia, HertzianaDP, from the Biblio-theca Hertziana collection, and SemArt+, refined from the SemArtdataset. In multimodal retrieval and art understanding, CANVAS out-performs the baselines, supporting the view that multi-relational align-ment is not just theoretically motivated but also practically essential.