Sector-Level Cross-View Geo-Localization with Implicit Orientation via Azimuthal Scanning
Abstract
Cross-View Geo-Localization (CVGL) is crucial for naviga-tion in GPS-restricted environments. However, bridging the severe per-spective gap between ground and aerial views remains challenging. Exist-ing approaches either rely on coarse image-level retrieval, suffer from un-stable optimization in pose regression, or require expensive fine-grainedannotations. These limitations suggest that current formulations of cross-view geo-localization are insufficient for reliable fine-grained localization.To address this limitation, we introduce a new task, Sector-Level Cross-View Geo-Localization (SLCVGL), which aims to perform fine-grainedsector-level geo-localization of a limited Field-of-View (FoV) ground im-age within an omnidirectional aerial image, while simultaneously infer-ring its ground-view orientation without manually annotated pose labels.To operationalize this task, we propose an Azimuthal Scanning frame-work. It decomposes aerial feature maps into multiple azimuthal sectors,instead of relying on holistic matching. By aligning ground features withthese sectors, the model explicitly retrieves a sector-level sub-region andimplicitly estimates orientation, eliminating the need for continuous poseregression supervision. This design also reduces interference from unob-served overhead regions. To benchmark the proposed task, we repur-pose standard panoramic datasets into a limited-FoV sector-level proto-col with random FoV cropping and soft labeling. Extensive experimentson CVUSA and CVACT demonstrate improvements over state-of-the-art methods, with a 30.17% R@1 improvement on CVACT under the 70°FoV setting, while maintaining strong robustness across varying FoVs.