Unsupervised Pixel-Level Semantic Left-Right Understanding of In-the-Wild Images
Abstract
While various works address reflective symmetry understand-ing in 3D data and images, pixel-level semantic left-right prediction ofin-the-wild images remains challenging, due to certain difficulties includ-ing the lack of 3D information, occlusion, object pose variation, partial-ity, etc. In this work, we propose an unsupervised learning framework totackle this challenge. Leveraging recent advances in vertex-wise semanticleft-right understanding of 3D data, our unsupervised learning methodjointly utilises 3D shape and image datasets to infer pixel-wise seman-tic left-right predictions in single-view images. In particular, we showthat a medium-scale 3D shape dataset comprising mainly of human- andquadruped animal-like shapes, combined with diverse in-the-wild imagedata, are sufficient to achieve high-quality semantic left-right predictionin images, even for entirely unseen 3D object categories, such as carsor trains. Overall, our approach achieves superior performance in densepixel-wise semantic left-right predictions on both rendered and in-the-wild image datasets when compared to existing state-of-the-art methods.