Skip to yearly menu bar
Skip to main content
Main Navigation
Select Year: (2026)
2026
2024
2022
My Stuff
Create Profile
Reset Password
Login
Getting Started
Schedule
Tutorials
Workshops
Main Conference
Keynotes and Panels
Orals
Spotlights
Papers
Paper Awards
Sponsors
Organizers
Help
Layout:
mini
compact
topic
detail
×
No topics available
No sessions available
title
author
topic
session
shuffle
by
serendipity
bookmarked first
visited first
not visited first
bookmarked but not visited
Loading...
Enable Javascript in your browser to see the papers page.
TooBad: Backdoor Diffusion Models with Ultra-Low Poison Rate and Imperceptible Trigger
From Local to Global: A Progressive Reconstruction Network for Diffractive Snapshot Spectral Imaging
NeSy-Route: A Neural-Symbolic Benchmark for Constrained Route Planning in Remote Sensing
GEO-Detective: Unveiling Location Privacy Risks in Images with LLM Agents
UMO: Unified In-Context Learning Unlocks Motion Foundation Model Priors
PHOSA: Photorealistic 3D Sign Avatar Modeling and Benchmark
GARDEN: Gravity-Aligned Reconstruction of Disentangled ENvironments from RGB images
EatVid-Bench: A Multimodal Fine-Grained Eating Behavior Video Dataset
Automatic Method Illustration Generation for AI Scientific Papers via Drawing Middleware Creation, Evolution, and Orchestration
AiSCREAM: Absolute Target Localization with Language-Conditioned Cross-View Alignment for Autonomous Vehicles
Towards More Efficient Decoding for Autoregressive Vision-language-action Models
Mechanistic interventions for explainable digital pathology uncovers adversarial vulnerabilities
GenLCA: 3D Diffusion for Full-Body Avatars from In-the-Wild Videos
SVI360: Spherical Video Interpolation
Fast-dVLA: Accelerating Discrete Diffusion VLA to Real-Time Performance
SeekFlow: Synergizing Radiology and Pathology Foundation Models for Precision Oncology via Knowledge-Guided Evidence Flow
DSH-Bench: A Difficulty- and Scenario-Aware Benchmark with Hierarchical Subject Taxonomy for Subject-Driven Text-to-Image Generation
IC-World: In-Context Generation for Shared World Modeling
Under One Sun: Multi-Object Generative Perception of Materials and Illumination
Generative Relightable Avatars
NoisEasier: Test-Time Noise Optimization for Text-to-Video Generation
Explicit Semantic–Spatial Alignment for Open-Vocabulary Object Detection
MotionSplicer: Part-Based Motion Editing for 4D Volumetric Videos
Less is More: Reducing Complexity in Vision-Language-Action Systems
Difficulty-Conditioned Attribute-Specific Restoration for Low-Light Image Enhancement
Frozen CLIP Priors for Robust Self-Supervised Poisson Inverse Problems
Online Segment 3D Gaussians via Launching Virtual Drones
RaysUp: Ultra-light Universal Feature Upsampling via Geometry-Aware Ray Representation
LoMa: Local Feature Matching Revisited
EGM: Efficient Visual Grounding Language Models
HAT-4D: Lifting Monocular Video for 4D Multi-Object Interactions via Human-Agent Collaboration
HairWeaver: Few-Shot Photorealistic Hair Motion Synthesis with Sim-to-Real Guided Video Diffusion
ZTRS: Zero-Human Demonstration End-to-end Autonomous Driving with Trajectory Scorer
Rolling Shutter Camera Self-Calibration
ContextFlow: In-Context Flow Matching for Robot Manipulation
Synesthesia via Direct Latent Augmentation: Bypassing the Decode-Encode Loop for Cross-Modal Distillation
SFDATrack: Generalized Source-Free Domain Adaptive Tracking Under Adverse Weather Conditions
OneHSI: A Unified Hyperspectral Foundation Model with Physical Consistency
Pro-Pose: Unpaired Full-Body Portrait Synthesis via Canonical UV Maps
Register Any Point: Scaling 3D Point Cloud Registration by Flow Matching
Hybrid Advantage Estimation with Unified Critic for VLM Agentic Reinforcement Learning
Understanding Geometric Representations in Self-Supervised Vision Transformers via Subspace Intervention
MemRoPE: Training-Free Infinite Video Generation via Evolving Memory Tokens
Global Graph-Validated Optimization for VLM-based 3D Indoor Scene Generation
R3RECON: Radiance-Field-Free Active Reconstruction via Renderability
InFlux++: Real and Synthetic Data for Estimating Dynamic Camera Intrinsics
OpenVE-3M: A Large-Scale High-Quality Dataset for Instruction-Based Video Editing
HCSU: A Dataset and Benchmark for Fine-Grained Historical Calligraphy Style Understanding
Region-Aware Multimodal Large Language Model via SlowFast Tokenization and Pseudo-Mask Guidance for 3D CT Report Generation
Odoriko: A Shape-Aware Multimodal Diffusion Framework for Human Motion
Invoice Haystack: Benchmarking Document Retrieval and Visual Question Answering Under Strong Visual Homogeneity
Foveated Reasoning: Stateful, Action-based Visual Focusing for Vision-Language Models
SAMPLe: A Sharpness Aware Minimization based Optimizer for Prompt Learning in Vision-Language Models
The Sterkfontein Caves Dataset: A Novel View Rendering Challenge from the Cradle of Humankind
Rdm: Re-conceptualizing Distribution Matching as a Reward for Diffusion Distillation
ConTrack: Constrained Hand Motion Tracking with Adaptive Trade-off Control
Rosetum3D: A Large-Scale 3D Vision Dataset from Preharvest Roses
MoCA3D: Monocular 3D Bounding Box Prediction in the Image Plane
On Locality and Length-Generalization in Visual Reasoning
SeeClear: Reliable Transparent Object Depth Estimation via Generative Opacification
Image Warping for Image-to-Image Translation
ClusterStyle: Modeling Intra-Style Diversity with Prototypical Clustering for Stylized Motion Generation
HuCollisionField: Resolving Self-Collisions via Neural Fields for Human Prediction
Urban Boundaries, Social Barriers: A Benchmark and Vision-Centric Framework for Mapping Gated Communities and Equity Implications
Free-Range Gaussians: Non-Grid-Aligned Generative 3D Gaussian Reconstruction
TaskTok: Delving into Task Tokens for Task-driven Image Restoration
Training-free Cross-domain Few-shot Segmentation via Robust Semantic Representation and Matching
Variational Patch Gating for Training-Free Few-Shot Classification
Vision-TTT: Efficient and Expressive Visual Representation Learning with Test-Time Training
TraversRL: Traversable Pedestrian Pathway Generation With Reinforcement Learning
Improved Immiscible Diffusion: Accelerating Diffusion Training by Reducing Miscibility
On the real-world generalisability of Optical Flow models
Entropy-Controlled Flow Matching
Beyond Final Answers: CRYSTAL Benchmark for Transparent Multimodal Reasoning Evaluation
TurboMPLE: Joint Infrared Turbulence Mitigation and Physical Fields Estimation via Mutual Progressive Layered Extraction
Learning Generatable Mutual Distance for Scene-Aware Human Motion Generation
2D Features Are All You Need for 3D Shape Understanding
Disentangling Hallucinations: Orthogonal Semantic Projection for Robust Interpretability
SRRA: Stable-Rank-Based Residual Adaptation for Generalizable Deepfake Detection
Articulat3D: Reconstructing Articulated Digital Twins From Monocular Videos with Geometric and Motion Constraints
Self-Evolving Agentic Image Restoration via Deliberate Planning and Intuitive Execution
FitControler: Toward Fit-Aware Virtual Try-On
Geometry-Aware Single-Image 4D Synthesis via Dense Trajectory Generation
On the Vulnerability of Parameter-Level Defenses to Model Merging
Improving Knowledge Distillation Under Unknown Covariate Shift Through Confidence-Guided Data Augmentation
Roam2Room: A Unified Floorplan-to-Furnished Framework for Controllable Indoor Scene Generation
Two-Parameter Flow Map Learning for Continuous-Time Diffeomorphic Image Registration
3D Scene-Adaptive Trajectory-Controllable Human Image Animation with Camera Movement
C3-Bench: A Context-Aware Change Captioning Benchmark
Ground3D-LMM: Fine-Grained 3D Point Grounding and Spatial Reasoning with LMM
Learning to Compose: Revisiting Proxy Task Design for Zero-Shot Composed Image Retrieval
Em-Garde: A Propose-Match Framework for Proactive Streaming Video Understanding
VISOR++ : VISUAL INPUT BASED STEERING FOR LARGE VISION LANGUAGE MODELS
WildProp: Visual Estimation of Wildlife Body Proportions at Scale
JSON: Jigsaw Self-play Optimization for Normalizing Flows
HiPolicy: Hierarchical Multi-Frequency Action Chunking for Policy Learning
PLOT: Pseudo-Labeling via Object Tracking for Monocular 3D Object Detection
Filterless Snapshot Hyperspectral Imaging using Guided Patch Diffusion
ATOMIC: A Domain-Specific Vision-Language Model for Transmission Electron Microscopy
Grounding World Simulation Models in a Real-World Metropolis
Mind-to-Face: Neural-Driven Photorealistic Avatar Synthesis via EEG Decoding
SAF3R: Dynamic Sparse Attention for Feed-Forward 3D Reconstruction Transformers
Hierarchical Spatial and Channel Aggregation for Cross-domain Few-shot Segmentation
PASDiff: Physics-Aware Semantic Guidance for Joint Real-world Low-Light Face Enhancement and Restoration
ReinDriveGen: Reinforcement Post-Training for Out-of-Distribution Driving Scene Generation
DARL: Efficient Document-to-Markup Generation via Look-Ahead Diffusion Trajectory Sampling
Why Do Vision Language Models Struggle To Recognize Human Emotions?
EgoMAN: Interaction-Structured Reasoning for Egocentric 3D Hand Trajectory Prediction
RPM-Distill: Physiology-guided Adaptive Cross-modal Distillation for Robust Remote Physiological Measurement
IConE: Batch Independent Collapse Prevention for Self-Supervised Representation Learning
Taming LLMs for Codematic Indoor Scene Generation
PMGC-SimVP: Parametric Multi-scale Gated Convolution for Global Ionospheric TEC Prediction
360° Image Perception with MLLMs: A Comprehensive Benchmark and a Training-Free Method
A Classifier-Agnostic Zero-Shot Adversarial Attack Detection via CLIP
BLOB-Q: Boosting Low Bit ViT Quantization via Global Optimization on Model Distortion
EmbodiedHead: Real-Time Listening and Speaking Avatar for Conversational Agents
From Reconstruction to Decision: A Post-Encoder Plug-in Adapter for Curvilinear Segmentation
Overlap-Consistent View Decomposition for Adapting Vision--Language Models to 360° Panoramas
FlowerDance: MeanFlow for Efficient and Refined 3D Dance Generation
Holo360D: A Large-Scale Real-World Dataset with Continuous Trajectories for Advancing Panoramic 3D Reconstruction and Beyond
LongVQUBench: Benchmarking Long-Term Video Quality Understanding of Vision-Language Models
Control-DINO: Feature Space Conditioning for Controllable Video Diffusion
CogSENet: Blind Image Deblurring with Blur-Conditioned Semantic Routing and Explicit Frequency Fusion
FaceArmor: A Universal Facial Image Protection Against Diffusion-Based Manipulations
IACD: Iterative Adversarial Collaborative Detection via Dual-Perspective Blind Spot Discovery
OpenGround: Planning-based Online Perception for Open-World 3D Visual Grounding
EmbedCopilot: Evaluating Vision-Language Models for Hardware-Aware Embedded System Development
Multi-Modal Controlled Coherent Motion Generation
ThermoGS: Decoupling Physical Surface Attributes for Spatio-Temporal Thermal Field Emulation via 4D Gaussian Splatting
AV2T-Gen: Aerial Visible to Thermal Generation with Environment and Vehicle State Guidance
MIMFlow: Integrating Masked Image Modeling with Normalizing Flows for End-to-End Image Generation
MCPNet:Masked Coordinate Pooling-based Attention Network for Medical Landmark Detection
Plug-and-Play Attention Linearization for Pretrained Transformers
UniDDT: Unifying Multimodal Understanding and Generation with Decoupled Diffusion Transformer
Dense Reward for Multi-View 3D Reasoning with Global Maps and Local Views
ProSR: Semantic-Prototype-Guided Discrete Modeling for Physically Consistent SAR Super-Resolution
SupIR-GS: Thermal Infrared Super-Resolution Novel View Synthesis with Imaging-Calibrated 3D Gaussian Splatting
TetraSDF: Analytic Isosurface Extraction with Multi-resolution Tetrahedral Grid
FoundationGeo: Learning Spatial Pixel-Wise Fields for Monocular Metric Geometry
EXPLORE-Bench: Egocentric Scene Prediction with Long-Horizon Reasoning
Stay Unique, Stay Efficient: Preserving Model Personality in Multi-Task Merging
Keep It Simple: Multi-Key Episodic Memory Retrieval for Ultra-Long Video Understanding
TriNLOS: Triplane Representations for Neural Non-Line-of-Sight Imaging
Skyfall-GS: Synthesizing Immersive 3D Urban Scenes from Satellite Imagery
EraseLoRA: MLLM-Driven Foreground Exclusion and Background Subtype Aggregation for Dataset-Free Object Removal
HO-Flow: Generalizable Hand-Object Interaction Generation with Latent Flow Matching
AgentVLN: Towards Agentic Vision-and-Language Navigation
TopoAgent: An Agentic Framework for Automated Topology Learning in Medical Imaging
Proto-Gaussian: MRI Modality Translation Based on Learnable Structural Prototypes and 2D Gaussian Splatting
Combining Discrepancy-Confusion Uncertainty and Calibration Diversity for Active Fine-Grained Image Classification
Test-Time Noise Guided Adaptation for Realistic Autoregressive Video Generation
Gravity-aware partially calibrated absolute pose estimation from affine- or rotation-covariant features
360CityArena: A Realistic Virtual Urban Navigation Benchmark for Embodied Agents
OmniFall: From Staged Through Synthetic to Wild, A Unified Multi-Domain Dataset for Robust Fall Detection
HoloTetSphere: Unified TetSphere Mesh Reconstruction for Physical Simulations
HVGCD:Rethinking Generalized Category Discovery through Hypothesis–Verification
Latent Fusion: Decoding Consolidated 3D Geometry from Feed-forward Geometry Transformer Latents
Auto3R: Automated 3D Reconstruction and Scanning via Data-driven Uncertainty Quantification
Mitigating Positional Leakage in 3D Masked Autoencoders for Robust Representation Learning
Diversity-Aware View Partitioning for Scalable VGGT
Holo-Captioning: A Comprehensive Textual View of 3D Scenes
HAD: Combining Hierarchical Diffusion with Metric-Decoupled RL for End-to-End Driving
Controllable Egocentric Video Generation via Occlusion-Aware Sparse 3D Hand Joints
Egocentric World Model for Photorealistic Hand Object Interaction Synthesis
WiFlow: Estimating Optical Flow using WiFi Channel State Information
EgoDyn-Bench: Evaluating Ego-Motion Understanding in Vision-Centric Foundation Models for Autonomous Driving
Beyond 2D Matching: A Unified Single-Stage Framework for Geometry-Aware Cross-View Object Geo-Localization
SFKD: Spatial–Frequency Joint-Aware Heterogeneous Knowledge Distillation via Multi-Level Wavelet Spectral Interaction
Spectral Consistent Flow for One-step 3D Medical Image Translation
Spectral Gradient Orthogonalization Improves Differentially Private Training at Scale
Preventing Expert Collapse in MoE-dVLMs via Modality-Wise Norm Alignment
Benchmarking MLLMs on Mistake Recognition and Explanation in Single-Step Components of Cooking
Bridge-UniPS: Bridging Calibrated Photometric Stereo toward Universal Photometric Stereo
PhysRVG: Physics-Aware Unified Reinforcement Learning for Video Generative Models
Learning to Generate Rigid Body Interactions with Video Diffusion Models
Poppy: Polarization-Based Plug-and-Play Guidance for Enhancing Surface Normal Estimation
Cube-Splat: High-Fidelity 360° Gaussian Splatting SLAM via Cubemap Factorization and Adjoint-Consistent Optimization
From Visual Primitives to Semantic Masks: Fine-Grained Visual-Linguistic Alignment for Open-Vocabulary Remote Sensing Image Segmentation
Geometry-Preserving in 3D Gaussian Splatting for LiDAR-Camera Extrinsic Calibration
SPARC: Scalable Path-Specific Counterfactual Fairness via Causal Conditional Independence
SIGMA-Lane: Scale-pyramId Gated MAmba for Temporally Consistent Video Lane Detection
Manifold-Aware Spectral Compaction: A Graph Signal Processing Perspective on Online Gaussian Reduction for 3DGS SLAM
Rethinking Temporal Modeling in Visual Object Tracking via Decoupled Auxiliary Supervision
FontCopilot: Towards Generalist Multimodal Large Language Models for Holistic Chinese Font Engineering
Blind to Position, Biased in Language: Probing Mid-Layer Representational Bias in Vision-Language Encoders for Zero-Shot Language-Grounded Spatial Understanding
DR-GS: Physically-Based Deformable and Relightable 2D Gaussians
RAGrasp: A Retrieval-Augmented Framework with Diversity-Aware Modeling for Dexterous Grasp Generation
Reasoning Path and Latent State Analysis for Multi-view Visual Spatial Reasoning: A Cognitive Science Perspective
Stealthy Multi-task Adversarial Attacks
SAEdit: Token-Level Control for Continuous Image Editing via Sparse Autoencoder
MuCHeR: Multi-Person Camera-Centric Human Detection, Mesh Recovery and Tracking
Think While You Map: Asynchronous Vision-Language Agents for Incremental 3D Scene Graphs
Pushing the Limits of High-Resolution Weather Forecasting through Data Scaling
Accurate Zero-shot Quantization via Hierarchical Teacher-Assistant Distillation
HiChor: Hierarchical Choreography Generation from Pop Music with Choreographic Primitives
Flexible Control of 3D CT Generation via Text and Semantically-Defined Segmentation Prompts
Partial Skeleton Visibility for Action Recognition: A Constrained Field-of-View Approach
EventVGGT: Exploring Cross-Modal Distillation for Consistent Event-based Depth Estimation
SafeGuard: A Multi-Agent Perception-Reasoning Framework for Social-Risk AI-Generated Video Detection
OCTOPUS: Multi‑Agentic Universal Compositional Visual Retrieval
CoSPlan: Corrective Sequential Planning via Scene Graph Incremental Updates
OmniPoser: Flexible Human Motion Recovery in the Wild with Masked Flow Matching
UHD-MFF: Shattering Barriers in Multi-Focus Ultra-High-Definition Image Fusion via Learnable Lookup Tables
SPARC: Single-Pass Scaling for Motion Forecasting with Conformal Bayesian Last Layers
MetricAnything: Scaling Metric Depth Pretraining with Noisy Heterogeneous Sources
PuzLM: Solving Jigsaw Puzzles with Sequence-to-Sequence Language Models
LogiCo: A Unified Framework for Logical and Structural Anomaly Detection
ARGENT: Adaptive Hierarchical Image-Text Representations
CameraAnything: Refilming Videos with Arbitrary Camera Control
UniTranslator: A Unified Multi-modal framework for End-to-end In-Image Machine Translation
Sound-based Multi-Person 3D Pose Estimation
ViewSpatial-Bench: Evaluating Multi-perspective Spatial Localization in Vision-Language Models
FlashBEV: Fast and Memory-Efficient Exact BEV Transformation with IO-Awareness
EchoStyle: Unlocking High-Fidelity Video Stylization with Reverse Data Synthesis
Mind2Cloud: EEG-to-Point Cloud Generation with Two-Granularity Diffusion Decoding
Video Can Teach PAN-Sharpening: PSF-Aware Cross-Domain Supervision
Infinite-Homography as Robust Conditioning for Camera-Controlled Video Generation
Raymap-Guided Coupling for Drift-Robust Unposed Feed-Forward 3D Reconstruction
DASAM3D: A Unified Foundation Model for Enhanced 3D Scene Reconstruction and Segmentation
UniTriSplat: A Unified 3D Gaussian Splatting Framework with Uniform Spherical Rasterization for Universal Cameras
Axolotl3D: a Unified Framework for Faithful 3D Shape Completion
Discrete Diffusion Models with MLLMs for Unified Medical Multimodal Generation
Stylized Video Generation via Decoupled Data Synthesis and Gated Style Token Injection
PADFormer: Pose-agnostic Anomaly Detection from Sparse View Images
On the Plasticity Collapse in Continual Machine Unlearning
FlowPainter: Inpainting Optical Flow via Confidence-Guided Completion
NURBS Splatting: A Unified Differentiable Rendering Framework for Vector Graphics
Self-supervised Garment Dynamics with Persistent Wrinkles
Lessons and Open Questions from a Unified Study of Camera-Trap Species Recognition Over Time
CHARTSTYLE-100K: A Large-Scale Dataset for Structured Visualization Style Transfer
JacobianAvatar: Temporally Consistent Semi-rigid Avatar Reconstruction from a Monocular Video
ScenarioControl: Vision-language Controllable Vectorized Latent Scenario Generation
Revisiting Avatar-As-Image: High-Fidelity Registration is All You Need
Improving Adversarial Robustness via Activation Amplification and Attenuation
Unwarping the Lens: A Physics-Grounded Approach to Video Glasses Removal
Delayed Bidirectional Alignment via Disentangled Audio Semantics for Audio-Visual Segmentation
PKINet-v2: Towards Powerful and Efficient Poly-Kernel Remote Sensing Object Detection
Gaussian Volumetric Representation for Efficient Shear–Warp Visualization
Long-term Traffic Simulation via Structured Autoregressive Modeling
GeoDetect: Geometric Adversarial Detection for VLPs
SiPhy: Single-Image Physical Property Reasoning
Unleashing the Power of Large-Scale ViT in Zero-Shot SBIR: A Strong Baseline with Multi-Layer Feature Aggregation
Monte Carlo Energy Aggregation for Mobile 3D Gaussian Splatting
RePer-360: Releasing Perspective Priors for 360° Depth Estimation via Self-Modulation
Clue Matters: Empower Video Reasoning with Brain-Inspired Latent Clue Learning
Dual Masked Generative Adversarial Transformer for Unsupervised Domain Adaptation
Describe-Then-Act: Proactive Agent Steering via Distilled Language-Action World Models
DiffuPrompt: Adapting Video Foundation Models to 3D Medical Volumes via Latent Trajectory Priors
LinStereo: Linear-Complexity Global Attention for Multi-Scale Iterative Stereo Matching
Predictive Photometric Uncertainty in Gaussian Splatting for Novel View Synthesis
ExpoMotion: A Large-Scale Benchmark and A Householder Projection Network for Multi-Exposure Fusion
Multi-THuMBS: Multi-person Tracking of 3D Human Meshes Beyond Video Shots
Track4World: Feedforward World-centric Dense 3D Tracking of All Pixels
MedQ-Deg: A Multidimensional Benchmark for Evaluating MLLMs Across Medical Image Quality Degradations
Event-driven Motion Deblurring via Trajectory-based Kernel Reconstruction
OCTA-SOT: Online Cross-Modal Trajectory Adjustment for RGBT Anti-UAV Single Object Tracking under Spatio-Temporal Misalignment
Controlling Motion Transfer in Diffusion Transformers via Attention Heads
Towards In-Context Tone Style Transfer with A Large-Scale Triplet Dataset
Decoupling Moment from Event for Video Temporal Grounding
ReSWD: ReSTIR‘d, not shaken. Combining Reservoir Sampling and Sliced Wasserstein Distance for Variance Reduction.
From Phase to Phenomenon: Self-Supervised Learning of Subsurface Scattering with Minimal Phase-shift Inputs
What You Ask is What You Ground: Bridging Question Intent to Temporal Evidence for Grounded VideoQA
UniStitch: Unifying Semantic and Geometric Features for Image Stitching
Wavelet-Driven Cross-Domain Consistency for Mixed-Supervised 3D Tumor Segmentation
VLZip: Unified Visual and Textual Compression for Interleaved Long-Context Modeling
Category-Level 3D Correspondence in Camera Space via Morphable Object Priors
ORFC: Orthogonal Reparameterization for Low-Bitrate ViT Feature Coding
CAST3D: Customizing Arbitrary 2D Assets into 3D World
SAM+D: Parameter-Efficient Dimensional Lifting of SAM-Family Models via Depth-Routed LoRA and Depth Shifting
Uncertainty-Driven Gaussian Sphere Propagation for 3D Semantic Segmentation
VC-VAE: Leveraging Video Codecs for Training-Efficient and High-Fidelity Video VAE
REDistill: Robust Estimator Distillation for Balancing Robustness and Efficiency
Scaling Laws for Black-box Adversarial Attacks
Learning to Deny: Action Denial in Multimodal Large Language Models
Rethink Backdoor Robustness in Vision Transformers
Beyond Alignment: A Generative Matching Paradigm via Flow Matching for Zero-Shot Skeleton-Based Action Recognition
PARL-VLA: Pruning-Aware On-Policy Reinforcement Learning for Vision-Language-Action Model
MuSViT: A Foundation Vision Model for Sheet Music Representation
Do Multimodal LLMs Understand Intraoral Dental Data? Dataset, Platform, and Baselines
SyncFix: Multi-View Consistent Diffusion Refinement of 3D Reconstructions
Predicting Consequences and Reinforcing Navigation Policies with Latent World Models
EgoExoMoCap: Distributed Human Motion Capture via Ego- and Exocentric Body Tracking from Head-Mounted Devices
SyncVL: Synchronizing Vision ⟷ Language Using Unsupervised Adaptation
From Blobs to Spokes: High-Fidelity Surface Reconstruction via Oriented Gaussians
SubSplat: High-Resolution Pixel-aligned 3DGS via Sub-pixel Gaussian Reparameterization
DiffusionVL: Translating Any Autoregressive Models into Diffusion Vision Language Models
Unsupervised Semantic Segmentation Facilitates Model Understanding
EffiDINO: Task-Specific Model Pruning via Gram Anchoring Subspace Consistency
VSDiffusion: Taming Ill-Posed Shadow Generation via Visibility-Constrained Diffusion
Multi-Head Normalization for Wide Vision Transformers
Delineating Knowledge Boundaries for Honest Large Vision-Language Models
Incremental Online Scene Reconstruction by 3D Gaussian Triangulation
Correlation-Weighted Multi-Reward Optimization for Compositional Generation
Don’t Starve the Boundaries: Boundary-Constrained Label Propagation for Weakly Supervised 3D Segmentation
SARA: Structure-Aware Riemannian-Guided Alignment for Drone Image-Text Retrieval
Leveraging Dark Knowledge for Intrinsic Multimodal Out-of-Distribution Detection
UniGP: Taming Diffusion Transformer for Prior-Preserved Unified Generation and Perception
AHOY! Animatable Humans under Occlusion from YouTube Videos with Gaussian Splatting and Video Diffusion Priors
MAC-Splat: Multi-Attribute Consistency for High-Fidelity Sparse-View Reconstruction
Mitigating Modality and Language-Style Gaps for Zero-Shot Video Moment Retrieval
Advancing WordArt-Oriented Scene Text Recognition: Datasets and Methods
VGGRPO: Towards World-Consistent Video Generation with 4D Latent Reward
MMDiff: Extending Diffusion Transformers for Multi-Modal Generation
Fisheye3R: Adapting Unified 3D Feed-Forward Foundation Models to Fisheye Lenses
RAGA: Real Time Ray Traced Gaussian Shadow Casting for 3DGS Avatar-Scene Interaction
Audio-Visual Camera Pose Estimation with Passive Scene Sounds and In-the-Wild Video
Semantic-Anchored Evidential Fusion for Domain-Robust Whole-Slide Survival Analysis
From Predictions to Embeddings: Dual Knowledge Distillation for Instance-Dependent Partial Label Learning
CerDETR: Cell-Prior Empowered DETR for Cervical Lesion Detection
Provable and Robust Wavefront Sensing via Self-Reference Interferometry
Inference-time Motion Calibration for Video Generation
DSAR: Dual-Stream Autoregressive Modeling of Temporal Cloth Dynamics for Photorealistic Animatable Avatars
A Dual-Transformer Architecture with Cross-Attention for Multi-Camera View Recommendation
From Sparse to Dense: Multi-View GRPO for Flow Models via Augmented Condition Space
Sparse-View Surface Reconstruction using Gaussian Splatting through High-Confidence Depth Propagation with Normal Priors
Seeing What Matters: Lesion-Aware High-Resolution Patch Discovery and Fusion for Chest X-ray Report Generation
PASTEL: Panoramic Alignment for Monocular 4D Scene Reconstruction
KineticGS: Momentum-driven Coherent 4D Gaussian Splatting for Monocular Dynamic Scene Reconstruction
GoStop: Reinforcement Learning for Adaptive Temporal Aggregation in Event-Based Feature Tracking
Steerable Visual Representations
Forecasting Animal Motion
Unified Panoramic–Gaussian Representation for Monocular 4D Scene Synthesis
Learn2Fold: Structured Origami Generation with World Model Planning
MultihopSpatial: Multi-hop Compositional Spatial Reasoning Benchmark for Vision-Language Model
Detecting Backdoors in Object Detection via Pre-NMS Prediction Distribution Shift
4DGS360: 360° Gaussian Reconstruction of Dynamic Objects from a Single Video
ComplexMimic: Human–Scene Interaction Imitation in Complex 3D Environments
SparseDriveV2: Scoring is All You Need for End-to-End Autonomous Driving
Swap the Right Identity: Spatio-Temporal Preference Optimization for Identity Swapping
Resolution-Agnostic Neural Operators for Multi-Rate Sparse-View CT
DiTex4D: Direct Text-Driven 4D Generation with Structured Latent Diffusion
PercepTax: Benchmarking Cross-Property Reasoning in Vision-Language Models
Incentivizing Vision Language Models to Search for Long Video Question Answering
EMOTE: Expressive Motion and Shape Disentanglement for Human Animation
HighlightBench: Benchmarking and Diagnosing Markup-Driven Table Reasoning in Scientific Documents
SONIC: Spectral Optimization of Noise for Inpainting with Consistency
ReInGS: Re-Initializing 3D Gaussians against Sparsity Discrepancy in Few-Shot Novel View Synthesis
Dynamic-Robust Photometric–Semantic Reconstruction for Open-Vocabulary 3D Scene Understanding
TDSR-VLA: Transition-aware Denoising Sequence Representations for Vision-Language-Action
BrepLLM: Enabling Large Language Models to Understand Boundary Representations
Robust and Efficient Monocular 3D Gaussian SLAM for Kilometer-Scale Outdoor Scenes
YesTrack: Referring Multi-Object Tracking via MLLM-based Yes/No Reasoning
GeCo: Evaluating Geometric Consistency for Video Generation via Motion and Structure
TRINITY: A Multi-Perspective Benchmark for Personal-Style Video Highlight Detection
Saber: Anchoring Semantics to Scale-Aware Kinetic Salience for Zero-Shot Skeleton Action Recognition
Adaptive Neural Dynamics for Robust Geometric LiDAR-Inertial State Estimation on UAVs
SignSparK: Efficient Multilingual Sign Language Production via Sparse Keyframe Learning
HolisticSemGes: Semantic Grounding of Holistic Co-Speech Gesture Generation with Contrastive Flow-Matching
High-speed Imaging through Turbulence with Event-based Light Fields
Equivariant Symmetry-Aware Head Pose Estimation for Fetal MRI
G2P: Gaussian-to-Point Attribute Alignment for Boundary-Aware 3D Segmentation
Metric-Bench: Exploring In-context Spatial Metric Reasoning in VLMs for Indoor Scenes
Towards Memory-Efficient Autoregressive Video Generation via Instance-Specific Parametric Absorption
TEASR: Training-Efficient Any-Step Diffusion Transformer for Real-World Image Super-Resolution
RASA: Disentangled Spatial-Motional Priors for Cross-Identity Character Animation
ECHO: Ego-centric Modeling of Human-Object Interactions
Self-Improving Diffusion Classifiers with Minority Preference Optimization
Triangular Consistency as a Universal Constraint for Learning Optical Flow
AnyView: Synthesizing Any Novel View in Dynamic Scenes
VQT: Vector Quantization Tuning for Efficient Fine-tuning and Compression of Pre-trained Vision Transformers
OccDirector: Language-Guided Behavior and Interaction Generation in 4D Occupancy Space
UniDriveDreamer: A Single-Stage Multimodal World Model for Autonomous Driving
ControlHair: Synergizing Physics Simulator and Video Diffusion for Controllable Dynamic Hair Rendering
SynFlow: Scaling Up LiDAR Scene Flow Estimation with Synthetic Data
Nexus-Vid: Efficient Frequency Bridging with Homogeneous Latent Space for Video Unified Models
ZeroSplat: Generalized Referring Segmentation in 3D Gaussian Splatting
4D-VGGT: A SpatioTemporal Foundation Model for Dynamic Scene Geometry Estimation
SFM: Taming State Space Models for Text-to-Motion via Spatial-Frequency Modeling
CLIMP: Contrastive Language-Image Mamba Pretraining
C3ASD: Multi-Level Consistency-Driven Representation Learning for Robust Active Speaker Detection
LaMP: Learning Vision-Language-Action Policies with 3D Scene Flow as Latent Motion Prior
MotionDreamer: Universal Skeletal Motion Generation for 3D Rigged Shapes
Geometric Context Transformer for Streaming 3D Reconstruction
SSBP: Stage-Specialized Block Pruning for Video Diffusion Models
DreamEdit3D: Personalization of Multi-View Diffusion Models for 3D Editing
Implicit Neural Representation for Spherical Harmonics Reconstruction of Motion-Corrupted Fetal Diffusion MRI
DIVA: Instruction-Aware Vision Token Pruning via Dual-Probe Attention Discrepancy
BlenderFusion: 3D-Grounded Visual Editing and Generative Compositing
CASA: Cross-Attention over Self-Attention for Efficient Vision-Language Fusion
DICE: Disentangled Instance-Class knowlEdge prompt tuning via SAE for Vision-Language Models
M2Tok: Multi-head Multi-codebook Discrete Action Tokenization for Vision-Language-Action Models
SynHMR: Synergistic Joint-Mesh Modeling for LiDAR-based Human Mesh Reconstruction
Weather-Conditioned Depth Anything
Fast Sam 3D Body: Accelerating SAM 3D Body for Real-Time Full-Body Human Mesh Recovery
Towards Real-World Wearable Motion Reconstruction
FlatLands: Generative Floormap Completion From a Single Egocentric View
Don’t Settle at the Mode! Mitigating Diversity Collapse in Pretrained Flow Models via Feature Self-Guidance
Vector Scaffolding: Inter-Scale Orchestration for Differentiable Image Vectorization
Precise Video-to-Audio Generation with Cross-Modal Alignment in Latent Space
DriftScope: Measuring The Hidden Effects of Diffusion Model Fine-Tuning
PhyDetEx: Detecting and Explaining the Physical Plausibility of T2V Models
Self-transcendence: Is External Feature Guidance Indispensable for Accelerating Diffusion Transformer Training?
TIIF-Bench: How Does Your T2I Model Follow Your Instructions?
STANCE: Controllable Video Generation for Structured Dynamics via Sparse-To-dense ANChored Encoding
SIFT: Self-Imagination Fine-Tuning for Physically Plausible Motion in Video Diffusion Models
HitMem: Hierarchical Temporal 3D Memory with Multi-Modal Context-Aware Retrieval for Dynamic Environments
XPos3R: Cross-Modal Transformer for Intraoperative 2D/3D Registration
SPEAR: A Simulator for Photorealistic Embodied AI Research
ScAle: Attention Head Scaling as a Minimal Adapter for Spatial Reasoning in Vision–Language Models
Out of Sight, Out of Mind? Evaluating State Evolution in Video World Models
Trust Your Instincts: Confidence-Driven Test-Time RL for Vision-Language-Action Models
Generalized Biomedicine Discovery
PWM-ArtGen: Part World Model for Articulated Object Generation
Human-like Object Grouping in Self-supervised Vision Transformers
Towards Geometry-Grounded Dense Semantic Matching with VGGT Priors
PolicyTrim: Boosting Intrinsic Policy Efficiency of Vision-Language-Action Models
Fourier Compressor: Frequency-Domain Visual Token Compression for Vision-Language Models
Compact Low-Cost Hyperspectral Imaging via Angular-to-Spectral Diversity Conversion
Structured Hyperedge Adaptation for Parameter-Efficient Fine-Tuning of Vision Transformers
DPGS: A Diffusion-Prior Guided Framework for Large-Scale 3D Gaussian Splatting Reconstruction
Reflect-R1: Evidence-Driven Reflection for Self-Correction in Long Video Understanding
Spectral Gating via Damped Oscillations for Adaptive Implicit Neural Representations
Activation Quantization of Vision Encoders Needs Prefixing Registers
MoGAN: Improving Motion Quality in Video Diffusion via Few-Step Motion Adversarial Post-Training
Progressive Pose-Guided 4D Animal Reconstruction from Monocular Video
Moving Beyond More Views: Redundancy-Aware Ego–Exo Fusion for Proficiency Estimation
CapFrame: Text-Instructed Viewpoint Grounding in 3D Gaussian Scenes via Geometric Pseudo Labels
Diffusion Models are Open-World Affordance Learners: Leveraging Generative Priors for 3D Affordance Learning
From One-to-One to Many-to-Many: Dynamic Cross-Layer Injection for Deep Vision-Language Fusion
RaPTGS: Render-Agnostic Post-Training Compression of 3D Gaussian Splatting
Mitigating Radar-Inertial Calibration Ambiguities via SO(3) Manifold Steering
Structure Gaussian Splatting SLAM
SMART: When is it Actually Worth Expanding a Speculative Tree?
Paying More Attention to Visual Tokens in Self-Evolving Large Multimodal Models
ARC-Loc: Leveraging Azimuthal Ray Convergence as a Geometric Cue for Direct Cross-View Localization
TIGER: Taming Identity, Geometry, and Generative Priors for High-Quality Face Video Restoration
VKSR: Scalable Kernel Surface Reconstruction Using Vecchia's Approximation
ActiveStructure: Plane Scene Graph-Guided Active 3D Gaussian Splatting
NoPA: Non-Parametric Online 3D Scene Graph Generation
TORA: Topological Representation Alignment for 3D Shape Assembly
Efficient Camera Pose Augmentation for View Generalization in Robotic Policy Learning
Your Data Manifold is Secretly a Reward Model: Shell-LCC for Text-to-Video Generation
TIDES: Time-Derivative Event Simulation via Deformable Reconstruction
InterPet4D: A Multimodal 4D Human-Pet Interaction Dataset for Pet Motion Generation
Geometric Gradient Rectification for Safe Open-Set Semi-Supervised Learning
Graph-GSReg: Leveraging 3D Scene Graphs for Gaussian Splatting Registration
Capacity-Controlled Multi-View Stylization of 3D Gaussian Splatting
PyraE2E: Enhancing End-to-End WSI Analysis via Cross-Scale Super-Resolution
Why Far Looks Up: Probing Spatial Representation in Vision-Language Models
VIEW2SPACE: Studying Multi-View Visual Reasoning from Sparse Observations
LEAP-VLA: Latent-Enhanced Action Prototyping via Continuous Residual Latent Spaces for Vision-Language-Action Models
TriFlow: Generating Artist-Like 3D Mesh Topology via Nearest-Vertex Vector Fields
MASS: Motion-Aligned Selective Scan for Flow-Based Video Frame Interpolation
EgoVITA: Learning to Plan and Verify for Egocentric Video Reasoning
CoLT: Teaching Multi-Modal Models to Think with Chain of Latent Thoughts
Learning Physics-based Forward Model Corrections in Unrolled Networks for Diffuser-based Imaging
Benchmarking Dynamic Affective Reasoning: A Viewer-Centric Video Emotion Dataset
MemoBench: Benchmarking World Modeling in Dynamically Changing Environments
GRADE: Benchmarking Discipline-Informed Reasoning in Image Editing
FedOT: Ownership Verification and Leakage Tracing via Watermarks for Federated LDMs
DiffPro: Joint Timestep and Layer-Wise Precision Optimization for Efficient Diffusion Inference
Understanding the Impact of Geometric Foundation Models on Vision-Language-Action Models
SyncLoop: A Multimodal Dual-Loop Framework for Self-Improving Mathematical Reasoning
LIBERO-Safety: A Comprehensive Benchmark for Physical and Semantic Safety in Vision-Language-Action Models
LoGAN: Multilingual Font Localization with Generative Agents
MV-GEL: Language-Driven Multi-View Geometric Entity Localization on Meshes
Accelerated Likelihood Maximization for Diffusion-based Versatile Content Generation
HIDA: A Human-Intuition-Guided Depth-Aware Framework for Zero-Shot Amodal Segmentation
ESC: Emotional Self-Correction for Reliable Vision-Language Models
MixTTA: Low-Rank Cross-Channel Mixing for Reliable Test-Time Adaptation
Recognizing Co-Speech Gestures in-the-Wild
Open-Weather Robust 3D Detection via Dual-Critic Diffusion Alignment
Instant Expressive Gaussian Head Avatars at Over 100 FPS
Calibrated Harmonic Overlaid Implicit Neural Representations for Multi-Dimensional Data
DecoupleGS: Interactive 3D Gaussian Splatting for End-to-End Autonomous Driving Testing
RCEdit-500K: Reference Completion for Image-Conditioned Image Editing
Task Alignment: A simple and effective proxy for model merging in computer vision
Do Vision Language Models Recognize Visual Ambiguity?
Attention is Case-Sensitive
Face Anything: 4D Face Reconstruction from Any Image Sequence
Versatile Editing of Video Content, Actions, and Dynamics without Training
AutoSpeed: Annotation-Free Stage-Adaptive Motion Speed Learning for Robot Manipulation
Circuit-MLLM: Topological Logic-Guided Latent-Space Visual Reasoning for Circuit Schematic Understanding
ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP
SelfMOTR: Revisiting MOTR with Self-Generating Detection Priors
TRiGS: Temporal Rigid-Body Motion for Scalable 4D Gaussian Splatting
UniPR-3D: Towards Universal Visual Place Recognition with Visual Geometry Grounded Transformer
Scene Generation at Absolute Scale: Utilizing Semantic and Geometric Guidance From Text for Accurate and Interpretable 3D Indoor Scene Generation
FoundDP: Revisiting Weak Disparity Observability in Dual-Pixel Depth Estimation
Looking Back and Forth: Cross-Image Attention Calibration and Attentive Preference Learning for Multi-Image Hallucination Mitigation
3D-Aware VLMs with Implicit and Explicit Geometries
λSplit: Self-Supervised Content-Aware Spectral Unmixing for Fluorescence Microscopy
OmniCamera: A Unified Framework for Multi-task Video Generation with Arbitrary Camera Control
TPCNet: A Low-Light Image Enhancement Network Inspired by Triple Physical Constraints
NGPS: Structure-Preserving Self-Supervised Denoising via Neighbor-Guided Patch Sampling
Code-in-the-Loop Forensics: Agentic Tool Use for Image Forgery Detection
VICAL: Vicinal Consistency Alignment for Long-Tailed Visual Recognition
Composing Driving Worlds through Disentangled Control for Adversarial Scenario Generation
Enhanced Neural Video Representation Compression with High Scalability
GenRecal: Generation after Recalibration from Large to Small Vision-Language Models
Topology-Weighted Effective Rank: A Zero-Cost Proxy for Training Dynamics Stability in Neural Architecture Search
Relaxed Rigidity with Ray-based Grouping for Dynamic Gaussian Splatting
Fidelity- and Perception-Aware Local Implicit Attention for Arbitrary-Scale Image Super-Resolution
DoCoG: Mask-based Multi-Type Grounded Chain-of-Thought for Document QA
Stochastic Optimal Control Sampling for Diffusion Inverse Problems
InfraNet: Quality-Aware RGB Guidance for Infrared Object Detection
REALM: An RGB and Event Aligned Latent Manifold for Cross-Modal Perception
Towards Long-Form Spatio-Temporal Video Grounding
PartCHOI: Part-Aware Guidance for Clothed Human-Object Interaction Generation
TimeWalker: Personalized Neural Space for Lifelong Head Avatars
Articulated Object Reconstruction from Rest-State Observation
Depth-guided Multi-view Exposure Bracketing for HDR Robot Vision
SAFE-Pruner: Semantic Attention–Guided Future-Aware Token Pruning for Efficient Vision-Language-Action Manipulation
AutoV: Loss-Oriented Ranking for Visual Prompt Retrieval in LVLMs
The Devil Is in the Dark Pixels: Toward Brightness Bias-Robust Denoising
RawGen: Learning Camera Raw Image Generation
SOMA: From Surface Observations to Muscle Anatomy
One-Shot Feed-Forward 360° Animatable Avatar via Inpainted UV-Space Gaussian Modeling
SPHERE: From MRI Sampling Mechanisms to Spatial Priors for Generalizable Brain Tumor Segmentation
AIMold: An Autonomous AI-based Pipeline for Complex Mold Design
Inference-Time Scaling of Diffusion Models via Progressive Pruning Search
3D Gaussian Splatting Compression with Object Scalability
CSS-BA: Gate Guided Column Space Search for Bundle Adjustment
Universal Image Immunization against Diffusion-based Image Editing via Semantic Injection
Asymmetric Anchoring: Opening the Black Box of MLLMs for Forgery Detection
UrbanAlign: Post-hoc Semantic Calibration for VLM-Human Preference Alignment
LaxMotion: Rethinking Supervision Granularity for 3D Human Motion Generation
Learning Zero-Shot Subject-Driven Video Generation Using 1% Compute
Tricam-rPPG: A Multimodal Multispectral Dataset for remote Photoplethysmography
RealGen: Photorealistic Text-to-Image Generation via Detector-Guided Rewards
Mapping Dark-Matter Clusters via Physics-Guided Diffusion Models
A second-order theory of texture for depth from focus
OmniPoint: Universal Monocular Metric Pointcloud from Any Camera
MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation
Spectral and Trajectory Regularization for Diffusion Transformer Super-Resolution
PRISM-VO: Scale-Aware Visual Odometry Using Photometric Plenoptic Bundle Adjustment
ARMS: Anchor–Relational Motion Streaming for Seamless Solo-Social Motion Transitions
ESNE: Efficient Surface Normal Estimation for LiDAR Point Clouds with Sequential Modeling and Variability Guidance
SFD-Net: Sharp Feature Detection Network Based on Local Geometric Features
Boosting 6D Object Pose Estimation via Monocular Depth Cues
Towards Robust Driving Perception: A Flexible Scale-Driven Family for Self-Supervised Monocular Depth Estimation
Perceiving Better Moments: Cover Frame Reselection and Enhancement for Live Photos with the Live2K Dataset
Moiré Video Authentication: A Physical Signature Against AI Video Generation
WildDepth: A Multimodal Dataset for 3D Wildlife Perception and Depth Estimation
HVA-Fusion:Hierarchical Velocity-Aware 4D Radar-LiDAR Fusion for Robust 3D Object Detection
From Macro to Micro: Benchmarking Microscopic Spatial Intelligence on Molecules via Vision-Language Models
Cross-View Yaw Estimation in Location Uncertainty with Line-Aligning Yaw Scoring
Infinite Gaze Generation for Videos with Autoregressive Diffusion
OmniRen: Neural Rendering wih Heterogeneous Scene Primitives
SceneHI: High-Resolution 3D-Consistent Scene Texturing with Controllable Illumination
PhyGDPO: Physics-Aware Groupwise Direct Preference Optimization for Physically Consistent Text-to-Video Generation
LiveEdit: Towards Real-Time Diffusion-Based Streaming Video Editing
Semantic Browsing: Controllable Diversity for Image Generation
DF3DV-1K: A Large-Scale Dataset and Benchmark for Distractor-Free Novel View Synthesis
GRF-Recon: Global Ray-Field Optimization for Long-Sequence Feed-forward Reconstruction
InceptionGS: Generative Bootstrapping for Large-Scale Gaussian Splatting under Unstructured View Sampling
Geometrically Consistent Multi-View Scene Generation from Freehand Sketches
Test-Time Registers as Global Priors for Tokenized Image Generation
FlexiBrain: Resolution-Agnostic Voxel-Level Encoding for Native fMRI
Open-Vocabulary BEV Segmentation with 3D-Aware Geometric Constraints
UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models
ReconPhys: Reconstruct Appearance and Physical Attributes from Single Video
GRAFT: Geometric Refinement and Fitting Transformer for Human Scene Reconstruction
MapDreamer: Aerial Imagery Conditioned Latent Diffusion For Lane Level Map Generation
CLDefocus: Physically Grounded Compound-Lens Defocus Blur Synthesis
DANTE-W: Diffuse Albedo Neural Texturing in the Wild
OmniX: Any-view and Any-time 4D reconstruction via Feed-forward Trajectory Fields
ResilPhase: Plug-and-Play Phase Mapping and Noise-Resilient Macro-Trajectory Extrapolation for Diffusion Acceleration
FedLAS: Feature-Modulated Bidirectional Label Smoothing for Neural Network Calibration
Kilometer-Vision: A New Frontier for Large-Scale Spatial Awareness in VLMs
G2FM: A Geodesic Flow Matching Framework with Geometric Prior for Category-Level 9-DoF Pose Estimation
VLA-R1: Enhancing Reasoning in Vision-Language-Action Models
TopoGS: Planar Reconstruction via Topology-Aware 3D Gaussian Splatting
NanoGS: Training-Free and Lightweight Gaussian Splat Simplification
Wat3R: Underwater 3D Geometry Learning without Underwater Annotations
Prosthesis-Aware 3D Human Pose Estimation: A Dataset and Benchmark for RSP Users
VoxAnchor: Explicit Voxel-Semantic Grounding for Spatial Understanding in Videos
TriMotion: Modality-Agnostic Camera Control for Video Generation
Lightweight Online Reinforcement Learning for Block Decomposition of CAD Models
RefracGS: Novel View Synthesis Through Refractive Water Surfaces with 3D Gaussian Ray Tracing
StructSplat: Generalizable 3D Gaussian Splatting from Uncalibrated Sparse Views
IRIS: Intersection-aware Ray-based Implicit Editable Scenes
CustomX: Unified Character, Action, and Scene Customization in Video World Models
World Reconstruction From Inconsistent Views
Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning
NaLA: A 3D Native LLM Layout Agent for High-quality 3D Scene Generation
GENA3D: Generative Amodal 3D Modeling by Bridging 2D Priors and 3D Coherence
VOID: Video Object and Interaction Deletion
Director: Instance-aware Gaussian Splatting for Dynamic Scene Modeling and Understanding
Geo-ID: Test-Time Geometric Consensus for Cross-View Consistent Intrinsics
AirSplat: Alignment and Rating for Robust Feed-Forward 3D Gaussian Splatting
Spectral-Aware Analytic Class-Incremental Learning for Long-Tailed Distributions
DIGS: Differentiable, Incremental, Global, Scalable Pruning for Language Models
FaCT-GS: Fast and Scalable CT Reconstruction with Gaussian Splatting
GAP-Track: Bridging the Resolution Gap for Cross-Resolution RGBT Tracking
FUSE: Filter-Free Unified Spatiotemporal Estimation of SpO2 via Wave-Transport Modeling
∂DIBR: Differentiable Depth Image-based Rendering for Fast Novel View Synthesis
VisTa3D: A Dataset and Benchmark for Vision, Tactile, and 3D Point Clouds-based Thin Object Reconstruction
GaINeR: Geometry-Aware Implicit Neural Representation for Image Editing
Rectified Embedding Flow Learning for Aerial Multi-view Geo-localization
Boosting 3D Foundation Models with Featureless Pose Optimization
Geometric Distillation from Rectified Stereo: Leveraging Epipolar Cues for Monocular Depth
PanoLess: Environment Reconstruction from Partial Reflective Views
FillGS: Filling Observation Gaps in 4D Gaussian Splatting via Viewpoint-Time Selection and Generative Refinement
iSyncTab: Learning Cross-Modal Feature Sequencing for Image-Tabular Data via Neural Synchrony
Making Avatars Interact: Towards Text-Driven Human-Object Interaction for Controllable Talking Avatars
Detect by Track: Making Detector-Free Matcher Trackable
Large-Scale High-Quality 3D Gaussian Head Reconstruction from Multi-View Captures
Confidence-Based Mesh Extraction from 3D Gaussians
FFAvatar: Feed-Forward 4D Head Avatar Reconstruction from Arbitrary Images
InstantHDR: Single-forward Gaussian Splatting for High Dynamic Range 3D Reconstruction
HASSL: Hierarchy-Aware Self-Supervised Learning Framework for Single Cell Microscopy
ROAR-3D: Routing Arbitrary Views for High-Fidelity 3D Generation
On-Policy Diffusion Reinforcement Learning Meets Off-Policy Quality Anchoring
V-HOLD: Stabilizing Flow Trajectories to Rethink the Edit–Preservation Trade-off
Edit in 2D, Verify in 3D: Reinforcement Learning for Multi-view Consistent Scene Editing
TripVVT: A Large-Scale Triplet Dataset and a Coarse-Mask Baseline for In-the-Wild Video Virtual Try-On
FPicker: Topology-Guided Evolution for Filament Tracing in Low-SNR Microscopy
DiaDem: Advancing Dialogue Descriptions in Audiovisual Video Captioning for Multimodal Large Language Models
Progressively Spiral Mamba Fusion for Multimodal Tracking
Bayesian Self-Attention with Local Pixel Correlations for Lightweight Denoising Transformers
Do Flat Minima Improve Sparse Novel View Synthesis?
Granular Semantic Cognition for Visible-Infrared Person Re-Identification
NeLU3D: Neural Inverse Structured Light without Modeling the Projector
BioMTBee: Biologically Constrained Multi-View Template-Based 3D Reconstruction of Bumblebee
NaP-Control: Navigating Diffusion Prior for Versatile and Fast Character Control
WilLaGS: Latent-Conditional 3D Appearance Fields for Robust Gaussian Splatting In-the-Wild
The Prism Hypothesis: Harmonizing Semantic and Pixel Representations via Unified Autoencoding
Map2World: Segment Map Conditioned Text to 3D World Generation
Learn to Rank: Visual Attribution by Learning Importance Ranking
RIGS: Radar-Informed Gaussian Splatting for Uncertainty-Aware 3D Occupancy and Motion Prediction
TecoPrompt: Temporal-Conservative Prompt Learning for Vision-Language Models
Consistent Monocular Depth Estimation with Contact Region Boundary-Aware Refinement
SafeSAE-VLA: Interpreting OpenVLA Progress Dynamics with Sparse Feature Analysis
Decoupled Illumination Priors for Spatially Controllable Multi-View Indoor Scene Relighting
SLAIR: Structured Latent Flow Matching for All-in-One Image Restoration
FUSE: A Flow-based Mapping Between Shapes
Doe-2: 3D Representation World Model for Unified Driving Scene Forecasting
OVGGT: O(1) Constant-Cost Streaming Visual Geometry Transformer
SkySplat-OV: Generalizable Language Gaussian Splatting for Open-Vocabulary Scene Understanding from Sparse Satellite Views
Coarse-to-fine Contrast: A Hybrid Self-supervised Method for Non-rigid 3D Shape Matching
Fast and Compact 3D Gaussian Splatting with Polarized Opacity Prior
Habitat-GS: A High-Fidelity Navigation Simulator with Dynamic Gaussian Splatting
EvoVLA: Self-Evolving Vision-Language-Action Model
CURE: Cumulative Knowledge Reuse for Efficient Device-Server Hybrid Inference in Vision-Language Models
Generalizable Neural Reconstruction of High-Fidelity Surfaces via Sparse Volumetric Representations
DINO-SLAM: DINO-Informed RGB-D SLAM for Neural Implicit and Explicit Representations
COSY: Compositional 3DGS Synthesis for Disentangled Human Head Editing
Rethinking Detection Calibration: A Coordinate Perspective
SkyLume: A Large-Scale Multi-Illumination Aerial Benchmark for Urban Scene Reconstruction and Beyond
CVSBench: A Comprehensive Benchmark for Cross-view Spatial Reasoning and Dreaming
MVGS: Multi-view Regulated Gaussian Splatting for Novel View Synthesis
Motion-aware Sparse Pipeline for Lightweight Object Tracking
Where to Look Matters: Learning Influential Views for VLM-based 3D Visual Grounding
LumiDepth: Stable Monocular Depth in Multi-Illumination Scenes
Neuromorphic X-ray Computed Tomography
NUN: Nested Unfolding Network for Real-World Concealed Object Segmentation
Revisiting the Volumetric Data of 4DME: Compression, Extension and Benchmarking for Micro-Expression Analysis
Fast and Flexible Robustness Certificates for Semantic Segmentation
Active View Selection with Perturbed Gaussian Ensemble for Tomographic Reconstruction
Hierarchical Prompt Injector for Domain Generalization Segmentation
Visual Spatial Tuning
LiteGS: a high-performance framework to train 3dgs in subminutes via system and algorithm codesign
MOCHA: Multi-modal Objects-aware Cross-arcHitecture Alignment
REFINE: Super-efficient Pruning for 3D Gaussian Splatting via Rendering-Free Primitive Importance
SemGAN: A Semantic and Hierarchical Adversarial Network for 3D Human Pose Estimation
Interact3D: Compositional 3D Generation of Interactive Objects
Know3D: Prompting 3D Generation with Knowledge from Vision-Language Models
When Specialists Meet Generalists: Segmenter-Coordinated Asymmetric Learning for Label-Deficient Concealed Object Segmentation
EmoteGPT: 3D Human Facial Expression from Natural Language Descriptions
WeatherReasonSeg: A Benchmark for Weather-Aware Reasoning Segmentation in Visual Language Models
HandSCS: Structural Coordinate Space for Animatable Hand Gaussian Splatting
AGE: Agentic Gaussian Editing in 3D Scenarios
Seeing Touch from Motion: A Unified Modality-Aware Visuo-Tactile Policy with Tactile Motion Correlation
HairOrbit: Multi-view Aware 3D Hair Modeling from Single Portraits
SKEL-CF: Coarse-to-Fine Biomechanical Skeleton and Surface Mesh Recovery
Posterior Samplings are Missing Modalities Generators for Medical Image Translation
Nexels: Neurally-Textured Surfels for Real-Time Novel View Synthesis with Sparse Primitives
MotionAnymesh: Physics-Grounded Articulation for Simulation-Ready Digital Twins
Forge4D: Feed-Forward 4D Human Reconstruction and Interpolation from Uncalibrated Sparse-View Videos
DeGuNet: Depth-Guided Ultra-Compact Backbones for Efficient LiDAR-Camera 3D Detection
Multi4D: High-Fidelity Dynamic Gaussian Splatting via Multi-Level Competitive Allocation
Gaussians on Fire: High-Frequency Reconstruction of Flames
EditVerse3D: High-Quality 3D Object Editing with Region-Aware Learning
Inclusive Interactive Collisions for Multi-View Consistent Compositional 3D Generation
StochasT: Learning with Stochastic Turn Depth for Visual Instruction Tuning
3DGS3: Joint Super Sampling and Frame Interpolation for Real-Time Large-Scale 3DGS Rendering
Learning Probabilistic Prompt for Continual Learning
FlexiAvatar: Unified 3D Gaussian Human Avatars Under Arbitrary Body Visibility
FreeGen: Feed-Forward Reconstruction–Generation Co-Training for Free-Viewpoint Driving Scene Synthesis
VIGS-SLAM: Visual Inertial Gaussian Splatting SLAM
Honey, I Shrunk the Arc de Triomphe!
MoCam: Unified Novel View Synthesis via Structured Denoising Dynamics
IMMoE: Incomplete Multi-View Anomaly Detection via Mixture of View Experts Fusion
Multi-view Multi-vehicle Driving Dataset for Novel View Synthesis
Enhancing Embodied Reasoning and Grounding by Novel View Synthesis
Last-Layer-Centric Feature Recombination: Unleashing 3D Geometric Knowledge in DINOv3 for Monocular Depth Estimation
PhysChoreo: Physics-Controllable Video Generation with Part-Aware Semantic Grounding
Ada-VNNs: Adaptive Equivariance for Vector Neural Networks
Edit3r: Instant 3D Scene Editing from Sparse Unposed Images
City-Level 3D Surface Reconstruction with Viewpoint Orientation Partitioning and Scene Completion
Video Generation Models Are Inherent Lighting Estimators
QVAM: Query-guided View-aware Adaptive Modulation for Aerial-Ground Person Re-Identification
Less is More: A Simple yet Effective Object-Centric Prompting Strategy for Vision-Language Reasoning in Autonomous Driving
Dynamic Inverse Rendering for Enhanced Material-Lighting Decomposition
MAVIN: Multi-Shot Audio-Visual Generation with Customized Narrative Control
PaD-GS: Leveraging Distortion Map for Panoramic Gaussian Splatting
OBBSeg: Irregular Lesion Segmentation under Oriented Bounding Box Annotations
StructureGS: Structure-aware Gaussian Splatting for Articulated Object Reconstruction
PhysPO: Physics-Aware Local Preference Optimization for Physically Consistent Video Diffusion
Geometry-Aware Style Transfer in 3D Gaussian Splatting
PriSplat: Propagating Reliable Multi-view Information for Distractor-Free 3DGS
EgoGVAE: Ego-body Mesh Reconstruction via Guided Variational Autoencoder
Denoising-GS: Gaussian Splatting with Spatial-aware Denoising
Tackling Misattribution in 3D Intrinsic Decomposition via Proximity Attention Point Rendering
Let's Reward Step-by-Step: Step-Aware Contrastive Alignment for Vision-Language Navigation in Continuous Environments
DefenseSplat: Enhancing the Robustness of 3D Gaussian Splatting via Frequency-Aware Filtering
Towards High-Resolution Visual Perception via Hierarchical Entity Exploration
DINOv3D: 2D-3D Joint Optimization for Unified Spatial Understanding
MindFlow: Harmonizing Cognitive Semantics and Acoustic Dynamics for Facial Animation Generation in Dyadic Conversations
ReSplat: Learning Recurrent Gaussian Splatting
BEV-GS: Feed-forward Gaussian Splatting in Bird’s-Eye-View for Road Reconstruction
Racing in Volume with Flow Ensembles
ReViV: Reconstructing the Viewer and the View in 4D from Monocular Egocentric Video
EAGS: Error-Aware Gaussian Splatting with Dual-Confidence-Guided Modeling for Uncalibrated Driving Scenes
LaVPR: Benchmarking Language and Vision for Place Recognition
Fast Spatial Memory with Scalable Elastic Test-Time Training
InteractiveAvatar: Real-Time Streaming Video Generation for Consistent and Intent-Aware Avatars
SA-ResGS: Self-Augmented Residual 3D Gaussian Splatting for Next Best View Selection
Weight-Space Mixture-of-Experts for Implicit Neural Representation Classification
IndoorSplat: Enhanced Indoor Scene Reconstruction with Structured 2D Gaussian Splatting
SignNet-1M: Large-Scale Multilingual Sign Language Video Dataset with Downstream Benchmarks
HiReFF: High-Resolution Feedforward Human Reconstruction from Uncalibrated Sparse-View Video
ReconDreamer-RL: Enhancing Reinforcement Learning via Diffusion-based Reconstruction
When 3D Gaussian Splatting Recovers Real Surfaces
PIC: Revisiting INR for Image Coding with Fast Encoding and Sub-Millisecond Decoding
Geometry-Propagated Gaussian Splatting for Aerial Sparse Novel View Synthesis
VLMSysTrojan: Stealthy System-Aware Backdoor Attacks Against Vision-Language Models
Neural Harmonic Textures for High-Quality Primitive Based Neural Reconstruction
StratoSplat: Taming Layered Regularities for Sparse Aerial 3D Gaussian Splatting
Accelerating Multimodal Large Language Models with Prior-Corrected Token Reduction
SMP-UWGS: Coupled Physics-Geometry Optimization for Scalable Multi-Partition Underwater 3D Reconstruction
TruthLens: Object Hallucination Detection via Self-Evaluating Truthfulness Scores in LVLMs
EditHF-1M: A Million-Scale Rich Human Preference Feedback for Image Editing
GKDT: General Keypoint Detection Transformer
Geometric Probing for Isotropic Optimization Manifold in Sparse-View 3D Gaussian Splatting
To Adapt or Not to Adapt? Selective Adaptation for Vision-Language Models
VersaViT: Enhancing MLLM Vision Backbones via Task-Guided Optimization
BLASt3R: Bundle Adjustment of Any Image Set with Multi-View Matching and Monocular priors
GeoMix: Descriptor-Free Visual Localization via Global Context and Multi-Detector Training
StreamGVE: Training-Free Video Editing via Few-Step Streaming Video Generation
Scaling Multi-Reference Image Generation with Dynamic Reward Optimization
Domain Generalization via Text-Anchored Information Bottleneck
Text-Guided 6D Object Pose Rearrangement via Closed-Loop VLM Agents
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models
Skin-R1: Clinical Knowledge-Guided Dermatological Diagnosis Using Vision-Language Models
See Only When Needed: Context-Aware Attention Intervention for Hallucination-Free LVLMs
DRIFT: Difficulty-aware Rectified Flows for Through-plane MRI Super-Resolution
UniTac: A Unified Multimodal Model for Cross-Sensor Tactile Understanding and Generation
Target-Bench: Can Video World Models Achieve Mapless Path Planning with Semantic Targets?
RSICCLLM: A Multimodal Large Language Model for Remote Sensing Image Change Captioning
Is Monitoring Enough? Strategic Agent Selection For Stealthy Attack in Multi-Agent Discussions
MobileVLA-R1: Reinforcing Vision-Language-Action for Mobile Robots
Exploring Efficient Reasoning Segmentation with Small Language Models
Following Motion for Sequential Modeling in Video Frame Interpolation
Beyond Language: Grounding Referring Expressions with Hand Pointing in Egocentric Vision
Beyond Atomic Layouts: Compositional Design Understanding with Vision-Language Models
Schroedinger’s Cat: Probabilistic Representation and Prediction of Potential Scene Kinematics
Accelerating Text-to-Video Generation with Calibrated Sparse Attention
Benchmarking Vision-Language Models for Microscopic Plant Image Understanding
Before Thinking, Learn to Decide: Proactive Routing for Efficient Visual Reasoning
HIVE: Understanding Post Hallucination Reasoning in Vision Language Models
VLA Knows Its Limits
VLA-Hijack: A Transferable Patch Attack against Vision-Language-Action Models via Visual Proprioception Hijacking
Neural Gate: Mitigating Privacy Risks in LVLMs via Neuron-Level Gradient Gating
Attention-based Vision-Language Memory for Spatial Reasoning
NaVLM-PVC: Progressive Visual Compression for Efficient Native-Resolution Encoding in MLLMs
E-TTS: A New Embodied Test-Time Scaling Framework for Robotic Manipulation
TARS: MinMax Token-Adaptive Preference Strategy for Hallucination Reduction in MLLMs
Gripper-aware Vision Language Action Models
MonoSR: Open-Vocabulary Spatial Reasoning on Monocular Images
RoboClaw: An Agentic Framework for Scalable Long-Horizon Robotic Tasks
DiCoBench: Benchmarking Multi-Image Fine-Grained Perception via Differential and Commonality Visual Cues
Context Blindness in DPO: Mitigating Object Hallucination in MLLMs via Context-Calibrated Preference Optimization
Steering 3D Generations: Preference Alignment via Direct Reward and Preference Optimization
PanoGrounder: Bridging 2D and 3D with Panoramic Scene Representations for VLM-based 3D Visual Grounding
Are Video Reasoning Models Ready to Go Outside?
ProactiveBench: Benchmarking Proactiveness in Multimodal Large Language Models
Concept-as-Tree: A Controllable Synthetic Data Framework Makes Stronger Personalized VLMs
Dual-Generalization-aware Minimization for Continual Fine-Tuning of Vision-Language Models
Gradient sparsity regularization for training unlearning-compatible models
VisWordBench: Bridging the Gap in Cross-modal Reasoning for Multimodal Large Language Models
The Cost of Reasoning: Chain-of-Thought Induces Overconfidence in Vision-Language Models
Towards Robustness against Typographic Attack with Training-free Concept Localization
Stokes-Informed Diffusion for Robust Linear Polarization Estimation
Anchored, Not Graded: How Vision-Language Models Fail at Slant-from-Texture Perception
Beyond Artifacts: Real-Centric Envelope Modeling for Reliable AI-Generated Image Detection
Kirin: Animal Motion Generation from In-the-Wild Video
MMBU: A Massive Multi-modal Biomedical Understanding Benchmark to Probe the Perception Capabilities of Vision-Language Models
Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models
HUGE-Bench: A Benchmark for High-Level UAV Vision-Language-Action Tasks
EgoEverything: A Benchmark for Human Behavior–Inspired Long-Context Egocentric Video Understanding in AR Environment
From Masks to Pixels and Meaning: A New Taxonomy, Benchmark and Metrics for VLM Image Tampering
Temporally Aware Densification for Dynamic 3D Gaussian Splatting
SPECSIA: Stylization Dataset for Novel-View Enhancement in Drawing-based 3D Animation
Direct Preference Optimization for Perceptual Alignment via Vision-Language Consistency
ME-IQA: Memory-Enhanced Image Quality Assessment via Re-Ranking
Reweighting Framewise Attention in Video Transformers for Facial Expression Understanding
VolSplat: Rethinking Feed-Forward 3D Gaussian Splatting with Voxel-Aligned Prediction
Transferability Between Understanding and Generation in Unified Multimodal Models
MinerU-Diffusion: Rethinking Document OCR as Inverse Rendering via Diffusion Decoding
Unified Multi-plane Autoregressive Diffusion for 3D Multi-Contrast MRI Synthesis
Ring Forcing: Towards Precise Long-Term Memory for Autoregressive Video Diffusion
AVSplat:Dense-View Feed-Forward 3D Gaussian Splatting with Assist-View Preconditioning
Frames2Residual: Spatiotemporal Decoupling for Self-Supervised Video Denoising
Denoising the Deep Sky: Physics-Based CCD Noise Formation for Astronomical Imaging
Do Not Leave a Gap: Hallucination-Free Object Concealment in Vision-Language Models
Exclusivity-Guided Mask Learning for Semi-Supervised Crowd Instance Segmentation and Counting
Anchor Forcing: Anchor Memory and Tri-Region RoPE for Interactive Streaming Video Diffusion
TransVLM: A Vision-Language Framework and Benchmark for Detecting Any Shot Transitions
MomentSeg: Moment-Centric Sampling for Enhanced Referring Video Object Segmentation
Semantic-Geometric Dual Compression: Training-Free Visual Token Reduction for Ultra-High-Resolution Remote Sensing Understanding
ShotStream: Streaming Multi-Shot Video Generation for Interactive Storytelling
Source-Agnostic Image Translation Based on Latent Aware Adaptive Masking
Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously
Linear Scaling Video VLMs for Long Video Understanding
GUIDE: Resolving Domain Bias in GUI Agents through Real-Time Web Video Retrieval and Plug-and-Play Annotation
GeoBrowse: A Geolocation Benchmark for Agentic Tool Use with Expert-Annotated Reasoning Traces
Fabric Image Demoiréing Benchmark from Synthesis to Restoration
STRIDE: When to Speak Meets Sequence Denoising for Streaming Video Understanding
LACON: Training Text-to-Image Model from Uncurated Data
Art Beyond Semantics: Sheaf-Informed Contrastive Learning for Multi-Relational Representations
PhysRAG: Enhancing Physics-Awareness in Video Generation via Retrieval-Augmented Generation
Objects as Audio-Visual Modal Sound Fields
ID-LoRA: Identity-Driven Audio-Video Personalization with In-Context LoRA
Representation Alignment for Just Image Transformers is not Easier than You Think
Learning Semantic-Robust Change Detection via Semantic-Invariant Self-Distillation
SOCO: Benchmarking Semantic Object Correspondence in Vision Foundation Models
Online Reasoning Video Object Segmentation
URHead: A Unified UV-Space Representation for Joint Mesh–3DGS Optimization in Head Avatars
Gen2Balance: Generative Balancing for Long-Tailed Video Action Recognition
Masked Depth Modeling for Spatial Perception
ProLaViT: Learning Progressive Latent Visual Thoughts in Structured Latent Space
GuideMe: Benchmarking Multi-Domain Task Guidance and Intervention in Streaming Video
Unlocking Few-Shot Capabilities in LVLMs via Prompt Conditioning and Head Selection
InterCMDM: Block-Causal Diffusion for Autoregressive Human Interaction Generation
MeanTalker: Efficient and Expressive Speech-Driven 3D Facial Animation via Geometric-Aware Mean Flow
ProtoMappingNet: Interpretable Hierarchical Prototypes through Relational Prototype Mappings
Delving into Latent Spectral Biasing of Video VAEs for Superior Diffusability
EgoPolice: A Benchmark for Egocentric Video Understanding in High-Stakes Police Body-Worn Camera Footage
SALT: Self-Consistent Distribution Matching with Cache-Aware Training for Few-Step Video Generation
OrthoTrack: Continuous 6-DoF UAV Trajectory Estimation Anchored in Public Orthophotos
ICLAgent: Integrated Circuit Footprint Geometry Labeling via LMM-empowered Multi-Agent Framework
Surprise Forcing: What to Remember, When to Skip in Long Video Generation
Pol-CACTI: A System and dataset forHigh-Speed Polarized Video Compressive Imaging
Stream3D-VLM: Online 3D Spatial Understanding with Incremental Geometry Priors
MagicPrompt: Ultra-Lightweight Prompt Tuning for Video Generation
One4D: Unified 4D Generation and Reconstruction via Decoupled LoRA Control
Two Birds, One Projection: Harmonizing Safety and Utility in LVLMs via Inference-time Feature Projection
Beyond Sequential Distance: Inter-Modal Distance Invariant Position Encoding
Test-time Counterfactual Calibration for Hallucination-Resistant Temporal Grounding
InterEdit: Navigating Text-Guided Multi-Human 3D Motion Editing
SwiftWA: An Efficient Action-Centered World-Action Model
PAI-Studio: Cinematic Video Background Replacement with Camera-Aware Motion
A Benchmark for Heterogeneous Stereo Deblurring with Physically- and Epipolar-constrained Cross Attention
MV-Forcing: Long Multi-View Video Generation via 4D-Grounded Spatio-Temporal Self-Forcing
SV-TAD: Native Sparse Convolutions for Efficient Temporal Action Detection
Better Call CineCrew: Consistent Ultra-Long Narrative-to-Film Generation
VectorReLoc: Reliable Vectorized SD Map Visual Re-localization with Contrastive Feature Alignment
Gaze-to-text Generation: Beyond Categorical Decoding of Human Attention
PhysDrape: Learning Explicit Forces and Collision Constraints for Physically Realistic Garment Draping
FastSTAR: Spatiotemporal Token Pruning for Efficient Autoregressive Video Synthesis
X-Stream: Benchmarking MLLMs as Multiplexers for Multi-Stream Understanding
HiAR: Efficient Autoregressive Long Video Generation via Hierarchical Denoising
ECoSim: Data Efficient Fine-Tuning for Controllable Traffic Simulation
Ink3D: Sculpting 3D Assets with Extremely Complex Textures via Video Generative Models
Horizon3D: Sparse Radar-Camera Fusion for Long-Range 3D Perception in Autonomous Driving
MultiMem: Measuring and Mitigating Memorization in Multi-Modal Contrastive Learning
Think While Watching: Online Streaming Segment-Level Memory for Multi-Turn Video Reasoning in Multimodal Large Language Models
When the Teacher Has More Bits: Self-Teacher Latent Distillation for Learned Image Compression
SpecEyes: Accelerating Agentic Multimodal LLM via Speculative Planning and Perception
EmbodiedVAE: Disentangled Video VAE for Efficient and Controllable Embodied Manipulation
S2Gest: Split-Scan State Space Models for Dynamic Hand Gesture Recognition
Latent Visual Diffusion Reasoning with Monte Carlo Tree Search
Beyond Random Sampling: Distribution-Aware Alignment for Semi-Supervised Medical Image Segmentation
Learning Egocentric Cues from Exocentric Video using Privileged Egocentric Supervision
GroundSet: A Cadastral-Grounded Dataset for Spatial Understanding with Vector Data
SyncCache: Exploiting Asymmetric Dynamics for Fast Audio-Driven Portrait Animation
DART: Deformable Adaptive Reasoning with Temporal Queries for Online Skeleton-Based Action Recognition
HOIMask: Towards Generative Masked Modeling for Human Object Interaction Generation
FingerCap: Fine-grained Finger-level Hand Motion Captioning
FreeSwim: Revisiting Sliding-Window Attention Mechanisms for Training-Free Ultra-High-Resolution Video Generation
Evaluating and Enhancing Negation Comprehension in Remote Sensing MLLMs
SpEmoC: A Balanced Speaker-Segment Multimodal Emotion Benchmark
Synthetic Visual Genome 2: Extracting Large-scale Spatio-Temporal Scene Graphs from Videos
LogFA: Efficient Feature-Space Data Augmentation for Egocentric Temporal Action Segmentation
Deformable Triangle Splatting: Flexible Primitives for Real-Time Radiance Field Rendering
OnPoint: Offline-to-Online Multi-Level Distillation for Point-Supervised Online Temporal Action Localization
DeforM: Reasoning-Guided Physics-Aware Video Generation via Spatial-Temporal Masking
SPAR: A Sequential Primacy and Attribution Ranking Framework for Skill Determination
VersatileMotion: A Unified Framework for Motion Synthesis and Comprehension
X2SAM: Any Segmentation in Images and Videos
OneVAE: Joint Discrete and Continuous Optimization Helps Discrete Video VAE Train Better
A Benchmark and Multi-Agent System for Instruction-driven Cinematic Video Compilation
FeatTracker: Short- and Long-Range Temporal Feature Consistency for Robust Underwater Object Tracking
Towards in-the-wild Egocentric 3D Hand-Object Pose Estimation
MMAgent-R2: Learning to Rerank and Reject for Agentic mRAG
Simple Filtering Improves Masked Autoencoders
Synthetic Sub-Aperture Phase Augmentation for Demosaicing 2×2 Shared Microlens Sensors
DeRA: Decoupled Representation Alignment for Video Tokenization
GryphOne: Symbol-Aware Masked Diffusion for Structural Refinement in Offline Handwritten Mathematical Expression Recognition
SHIFT: Motion Alignment in Video Diffusion Models with Adversarial Hybrid Fine-Tuning
ClearText-Video: A Large-Scale Text-Centric Video Dataset Bridging Video Restoration and Scene-Text Enhancement
Sparse auto-regressive modeling for scene generation from multi-view images
DeCo: Zero-Shot Anomaly Generation through Decoupling and Recoupling
FLEG: Feed-Forward Language Embedded Gaussian Splatting from Any Views via Compact Semantic Representation
LoT-Pass: Long-term-robust Image Watermarking for Image to Video Generation
TinyHistory: Lightweight Video History Embeddings via Two-Stage Context Learning
ChronusOmni: Improving Time Awareness of Omni-Modal Large Language Models
Autoregressive Image Generation Needs Only a Few Lines of Cached Tokens
Audio-Visual Continual Test-Time Adaptation without Forgetting
RoboMirror: Understand Before You Imitate for Video to Humanoid Locomotion
LoCA: Spatially-Aware Low-Rank Convolutional Adaptation of Vision Foundation Models
Structural Assessment for Understanding and Guiding Dataset Distillation in Discrete Token Space
Adapting MLLMs for Nuanced Video Retrieval
TAR: Temporal Anchor-Constrained Reasoning for Video Temporal Grounding
Adaptive Latent Trajectory Anchoring for Action Segmentation Dataset Condensation
SharpGS: Sharpness-Preserving 3D Gaussian Splatting with Differentiable Blur-Driven Density Control
Memory-V2V: Memory-Augmented Video-to-Video Diffusion for Consistent Multi-Turn Editing
Safe Responses Matter: Output-Aware Safety Guardrail Mitigate Over-Refusal in MLLMs
HomeGuard: VLM-based Embodied Safeguard for Identifying Contextual Risk in Household Task
Pistachio: Towards Synthetic, Balanced, and Long-Form Video Anomaly Benchmarks
Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance
VoCa: Unified Autoregressive Modeling for Talking Audio-Video Generation
PanoSAM2: Lightweight Distortion- and Memory-aware Adaptions of SAM2 for 360 Video Object Segmentation
Video-Holmes: Can MLLM Think like Holmes for Complex Video Reasoning?
LOOM: Weaving Geometry-Consistent Human-Object Interaction Videos via Progressive Curriculum Learning
GimbalDiffusion: Gravity-Aware Camera Control for Video Generation
Guiding the Blind: Generalizing GUI Agents to Unseen Websites via Multimodal Tutorials
Test Time Training for Long Videos via Frame Forgetting Network
Trajectory-Level Continuous Action Representation for Robotic Manipulation
SIGNER: Temporally Grounded Sign Language Generation via Time-Resolved Conditioning
AffoGato: Open-Vocabulary Affordance Grounding with Automated Data Generation at Scale
PhenoLIP: Phenotype Guided Medical Vision–Language Pretraining
SPAR: Semantic-Pixel Self-Alignment and Adaptive Routing for Unified Multimodal Models
Towards Benign Memory Forgetting for Selective Multimodal Large Language Model Unlearning
From Evaluation to Enhancement: Benchmarking and Improving Think-with-Video Reasoning for Video Generative Models
SAM-MT: Real-Time Interactive Multi-Target Video Segmentation
Robust 3DGS-based SLAM via Adaptive Kernel Smoothing
V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning
EgoSim: Egocentric World Simulator for Embodiment Interaction Generation
CritiqueDriveVLM: From Verifier-Guided Reinforcement Learning to Latent Thought Distillation for Autonomous Driving
DG-Force: Disentangling and Gathering Forensic Cues is Needed for Image Manipulation Localization
Conditional Flow Matching for Visually-Guided Acoustic Highlighting
Geometry Grounding: Elevating Blind Distortion Correction with 3D Structural Priors
Triangle Splatting SLAM
PeCA: Palette Context Assisted Inference for Test-Time Paint-Bucket Colourisation on Animation Videos
From Passive Observer to Active Critic: Reinforcement Learning Elicits Process Reasoning for Robotic Manipulation
AnaPFL: When Closed-Form Solutions Meet Generalizationand Personalization in Personalized Federated Learning
Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation
SDSA: Shallow-Deep Squeezing Adapter for Vision-Language Models
Counterfactual World Models via Digital Twin-conditioned Video Diffusion
VIGOR: VIdeo Geometry-Oriented Reward for Temporal Generative Alignment
Scaling Verification Can Be More Effective than Scaling Policy Learning for Vision-Language-Action Alignment
MG-RWKV: Multi-Grained Context-Aware RWKV for Temporal Forgery Localization
ProAct: Agentic Lookahead in Interactive Environments
Whence the Voice? Self-supervised Dual-source Audio-Visual Localisation via Selective Convergence
WorldWander: Bridging Egocentric and Exocentric Worlds in Video Generation
Towards Temporal Compositional Reasoning in Long-Form Sports Videos
Single-Query Person-Centric Bimanual Hand-Object Interaction Detection
Pathwise Test-Time Correction for Autoregressive Long Video Generation
What Matters in RL-Based Methods for Object-Goal Navigation? An Empirical Study and A Unified Framework
DLGStream: Dynamic Language-embedded Guassian Splatting for Open-vocabulary Enabled Free-viewpoint Video Streaming
Structured Redundancy Modeling for Efficient Visual Token Pruning in High-Resolution MLLMs
VideoSfM: Exploiting Temporal Structure for Video-Based Structure-from-Motion
StrucTab: A Structured Optimization Framework for Table Parsing
VideoTIR: Accurate Understanding for Long Videos with Efficient Tool-Integrated Reasoning
SVI-Bench: A Dynamic Microworld for Strategic Video Intelligence
Natural Image Pretraining Improves Abstract Reasoning
AutoCompass: Accurate Visual Localization on Public Maps by Learning from Weak Labels
Hyperbolic Hierarchical Clustering for Visual Representation Learning
SegDiff: Segmented Trajectory Diffusion for Consistent and Adaptive Robot Manipulation
JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing
STAC: Selective Spatiotemporal Aggregation and Compression for Video Reasoning Segmentation
Text Dictates, Music Decorates: Energy-based Attention for Editable Dance Generation
GeoNVS: Geometry Grounded Video Diffusion for Novel View Synthesis
CoDePose: Multi-View 3D Human Pose Estimation via Coupled 2D-3D Denoising Diffusion
Ego-Human Motion Prediction with 3D-Aware LLM
SENTRY: SAM2-Enhanced Neighbor-Aware and Temporally Reasoned Memory for Visual Tracking
GAINS: Gaussian-based Inverse Rendering from Sparse Multi-View Captures
Token-Based Affordance Grounding with Large Vision-Language Models
DualCount: Structurally Consistent Density and Point Modeling for Zero-Shot Object Counting
OmniFace: Bridging the Image-to-Video Gap for High-Fidelity Face Swapping via Diffusion Transformer
Sticking Information in Plain Sight: Encoding and Detecting Hidden Stickers in the Real World
Learning to Suppress SPAD-based LiDAR Flare
Q-TriM: Question-Guided Tri-Modal Attention for Audio–Visual Question Answering
LINA: Learning INterventions Adaptively for Physical Alignment and Counterfactual Generation in Diffusion Models
CoIn: Comprehensive 2D-3D Inpainting with Gaussian Splatting Guidance
Demystifing Video Reasoning
Table-MCR2TR: Merged-Cell-Aware Table Recognition via Reinforced Multimodal Language Models
Recurrent Autoregressive Diffusion: Global Memory Meets Local Attention
Seeing Fast and Slow: Learning the Flow of Time in Videos
LENS: Adaptive Spatio-Temporal Zooming for Keyframe Sampling in Long-Form Videos
Open-Vocabulary Long Term Action Anticipation
MGM-Omni: Scaling Omni LLMs to Personalized Long-Horizon Speech
DASH: Dynamic Audio-Driven Semantic Chunking for Efficient Omnimodal Token Compression
Thinking in Streaming Video
Beyond Time Shifts: Adapting Omni-LLM as a Reference-Free Evaluator for Generative Audio-Visual Models
Disentangling Rotation and Translation from SE(3)-Equivariant Features for Shape Assembly
TaxoGrasp: Taxonomy-Guided Human Grasp Synthesis with Sparse Contact Constraint
Video-HolmesV2: Can MLLMs Reason with Spatio-Temporal Audio-Visual Evidence in Long Videos?
Conversational Human Audio-visual Talking Dialogue Generation
SemCityLoc: Aerial 6DoF Localization Using Semantic 3D City Models
SIGNET: Motion-Level Knowledge Transfer for Cross-Language Sign Language Translation
Reward Modeling for Computer-Using Agent from Video Execution
SignBind-LLM: Multi-Stage Modality Fusion for Sign Language Translation
PerceptionComp: A Video Benchmark for Complex Perception-Centric Reasoning
EgoCogNav: Cognition-aware Human Egocentric Navigation
STVFocus: Query-guided Spatio-Temporal Visual Focusing for Video LLMs
LARY: A Latent Action Representation Yielding Benchmark
NearID: Identity Representation Learning via Near-identity Distractors
Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models
From Script to Shot: A Benchmark for Grounding Screenplays in Movies
Sim, Yet Same: Physics-Aligned Simulator as Zero-Shot Data Scaler in Deformable Worlds
GTR: Guide-Then-Refine Token Compression for Training-Free Acceleration of Video-LLMs
Distribution-Alignment Bridge for Uncertainty-Aware Text-to-Video Retrieval
Decompose, Compare, and Decide: Multimodal LLMs are Implicit Few-Shot Learners
RoboGesture: Real-Time Semantic-aligned Co-Speech Gestures Generation for Humanoid Interaction
Keep-or-Drop? Adaptive Tokenizer for Compact Video Representation
StyleFusion360: View-Consistent Head Stylization via Adaptive Style Modulation
S3-Prune: Stability-Aware Token Budgeting for Long-Form Video-Language Models
Beyond Script Family Boundaries: Towards Unified Open-Set Scene Text Recognition
Layer-Aware Video Composition via Split-then-Merge
H2SVC: Head-aware Heterogeneous Streaming Video Cache for Online Video Understanding
Narrative-Driven Paper-to-Slide Generation via ArcDeck
EgoPHI: Estimating 3D Hand-Object Contact and Force from Egocentric Vision
CORE-V: Chain-Of-thought REasoning for Image Editing with Visual Interaction
PriSM: Parsing and Style-Mixed Consistency for Unsupervised Domain Adaptation in Facial Landmark Detection
ActionPlan: Future-Aware Streaming Motion Synthesis via Frame-Level Action Planning
Representations Before Pixels: Semantics-Guided Hierarchical Video Prediction
No Place to Hide: Benchmarking Video Hallucination with Background-Controlled Pairs
Accelerating Diffusion Models via Equal-Risk Caching
Continuous Heart Rate Variability Estimation from Egocentric Systems for Skill Assessment
AVTok: 1D Unified Tokenization for Holistic Audio-Video Generation
MoVA: Learning Asymmetric Dual Projections for Modular Long Video-Text Alignment
DiT as Real-Time Rerenderer: Streaming Video Stylization with Autoregressive Diffusion Transformer
UniScale: Arbitrary-Scale Anomaly Generation
Learning Consistent Temporal Grounding between Related Tasks in Sports Coaching
GEAR-Seg: A Grounded Explainable Agent for Reasoning Segmentation and Data Engine
StoryBlender: Inter-Shot Consistent and Editable 3D Storyboard with Spatial-temporal Dynamics
Natural Language Camera Movement Understanding
UniTemp: Unlocking Video Generation in Any Temporal Order via Autoregressive Distillation
JAM-Flow: Joint Audio-Motion Synthesis with Flow Matching
Fine-Grained Text-to-Video Retrieval for Camera-Trap Data
MJEPA: A Simple and Scalable Joint-Embedding Predictive Architecture for Audio-Visual Learning
GRE-Diff: Gaussian Room Embeddings for Structured Layout Diffusion
Benchmarking Scientific Understanding and Reasoning for Video Generation using VideoScience-Bench
Causal Yet Future-Aware: Dual-Path Temporal Modeling for Online Action Segmentation
DiTailed: Ensuring Visual Object Consistency in Text-Image-to-Image Flow Matching Models
EventMemAgent: Hierarchical Event-Centric Memory for Online Video Understanding with Adaptive Tool Use
Small Vision-Language Models are Smart Compressors for Long Video Understanding
FeVOS: Foresight Expression Video Object Segmentation
Histopathology Multi-modal Embedding for Pathology Composed Retrieval
Unified Video Dense Prediction from Disjoint Data
Explicit Logic Channel for Validation and Enhancement of MLLMs on Zero-Shot Tasks
Learning to Mask: Cross-Modal Noise Modulation for Hallucination Mitigation in Multi-modal Large Language Models
FEEL (Force-Enhanced Egocentric Learning): A Dataset for Physical Action Understanding
Unifying CNNs and ViTs for Learning-Efficient and Scalable Variational AutoEncoder
RotateAttention : RoPE-Aware Rotation and Range Rectification for INT4 Quantized Attention in Video Generation
CTEPM: Continuous-Time Event Process Memory for Long-Video Language Models
Self-Evolving MCP-GUI Agents via Automated Environment Generation and Experience Learning
LaViT: Aligning Latent Visual Thoughts for Multi-modal Reasoning
QSVideo: Query-Conditioned Semantic Temporal Retrieval for Video Understanding
ReactVAU: A Slow-Fast Decoupled Framework for Streaming Video Anomaly Understanding
Neural Collapse-Inspired Multi-Label Federated Learning under Label-Distribution Skew
Remembering Across Blocks: Topology-Conditioned Block-Progressive Memory for Skeleton-Based Action Recognition
HAS: Highlight-guided Attention Steering for Multimodal LLM Video Summarization
NanoVSR: Towards Real-Time Video Super-Resolution on Edge Devices
Reasoning-Guided Part-Level Visual Grounding via Reinforcement Learning
LogicIR: Logic Gate Networks for Image Restoration
Freqformer: Image-Demoiréing Transformer via Effective Frequency Decomposition
NeuroRefiner: Morphology-Aware Multi-Agent Refinement for 3D Fluorescence Microscopy Neuron Segmentation
Spatiotemporal Flux Probing for Single-Photon Videography
Atlas is Your Perfect Context: One-Shot Customization for Generalizable Foundational Medical Image Segmentation
Raw-JPEG Adapter: Efficient Raw Image Compression with JPEG
LiDAR-EVS: Enhance Extrapolated View Synthesis for 3D Gaussian Splatting with Pseudo-LiDAR Supervision
Gaussian Belief Propagation Network for Depth Completion
ECHO: Efficient Chest X-ray Report Generation with One-step Block Diffusion
WAFT-Stereo: Warping-Alone Field Transforms for Stereo Matching
MambaRaw: Selective State Space Modeling for Efficient 4K RAW Image Reconstruction
AlphaRad: Grounded Zero-Shot Classification in Chest Radiology via α-Corrected Binary Cross Entropy and Factorized Latent Supervision
Memory-Supported Synergistic Adaptation for Training-Free Test-Time Medical Image Segmentation
IDeaL: Data-Free Multi-Teacher Distillation via Improved Dead Leaves
Mimicking Radiologists: A Coarse-to-Fine Framework with Structural Sparse Tokens for Dual-LLM Computed Tomography Report Generation
Kiroshi: An Agentic Perception System for High-Accuracy Image Parsing
EcoVideo: Entropy-Orchestrated Video Generation Paradigm in Cloud-Edge Dynamics
Hybrid Event–Frame Sensors: Modeling, Calibration, and Simulation
MedSynapse-V: Bridging Visual Perception and Clinical Intuition via Latent Memory Evolution
VesselTok: Tokenizing Vessel-like 3D Biomedical Graph Representations for Reconstruction and Generation
Quantile‑Adaptive Temperature Scaling for Confidence Calibration
Hyper-Network Neural Functional Maps for Unsupervised Robust 3D Shape Matching
MediRound: Multi-Round Entity-Level Reasoning Segmentation in Medical Images
Differentiable Polarized Path Tracing
If It's Not Efficient, It's Not Usable: Real-Time OOD Detection with Latent De-Biasing and High-Quality Negative Samples
Masked BRep Autoencoder via Hierarchical Graph Transformer
SparseCtrl-HOI: Sparse Temporal Control for Human-Object Interaction Video Generation
RainODE: Continuous-Time Precipitation Forecasting with Latent Neural ODEs
MVFusion-GS: Motion-Variance Guided Temporal Attention for High-Quality Dynamic Gaussian Splatting
TopoGAT: Plug-and-Play Topological Graph Attention for Fine-Grained 3D Segmentation
Learn to See the Unseen in Low-light Spike Streams
2K Retrofit: Entropy-Guided Efficient Sparse Refinement for High-Resolution 3D Geometry Prediction
BWAFDA: Block-wise Weighted Attention Fusion with Detail-aware for No-Reference Image Quality Assessment
Concept-to-Pixel: Prompt-Free Universal Medical Image Segmentation
CUST : Clustered Unit-level Similarity Transformer for Lightweight Image Super-Resolution
DeLux: Cross-Modal Local Artifact Restoration in Video Using Neuromorphic Data
EvDiff: High Quality Video with an Event Camera
Capacity Overflow: A Blind Spot for Backdoor Attacks in Vision MoE
Dual-Output Multi-Exposure HDR Reconstruction via SDR Fusion and Gain Map Inverse Tone Mapping
Multi-Channel Uncertainty-Weighted Score Matching for Conditional Diffusion in Medical UDA
UniH3: Unifying Hierarchical Homogeneity and Heterogeneity for All-in-One Medical Image Restoration
MedCAGD: Context-Aware Gated Decoder for Robust Medical Image Segmentation
Foundation-Guided Representation Alignment for Multimodal Medical Image Registration
Φeat: Physically-Grounded Material Feature Representation
From Multi-Resolution Cells to Gigapixel Whole Slide Images Foundation Model for Computational Pathology
Progression as Latent Drift: Generative Forecasting of Slow-Evolving Pathologies
SDUM: A Scalable Deep Unrolled Model for Universal Cardiac MRI Reconstruction
Lost in the Tail: Addressing Geographic Imbalance in Urban Visual Place Recognition
Towards Trustworthy Dermatology MLLMs: A Benchmark and Multimodal Evaluator for Diagnostic Narratives
Physics-Grounded Disentangled Flow Modeling for Brain Disease Progression Trajectory
ConfCtrl: Enabling Precise Camera Control in Video Diffusion via Confidence-Aware Interpolation
From Minimal Clinical Prompts to 3D: Spacing-Aware Prompt Propagation for Multimodal Prostate Lesion Segmentation in bpMRI
340 FPS Reflection-free Video from Spikes Modulated by a Rapidly Rotating Polarizer
Clinical Cognition Alignment for Gastrointestinal Diagnosis with Multimodal LLMs
Score-Based Matching with Target Guidance for Cryo-EM Denoising
Event-based Sparse-view Background-Oriented Schlieren Tomography
h-Flow: Flexible Flow-based Image Editing via Doob's h-Transform
Multi-Block-Attention-based Color Constancy
G-ZAP: A Generalizable Zero-Shot Framework for Arbitrary-Scale Pansharpening
Proximity-Constrained Counterfactual Decoding for Hallucination-Robust Medical VQA
DiffVP:Differential Visual Semantic Prompting for LLM-Based CT Report Generation
Finding Highlight Images In Your Albums:From Benchmark To MLLM
MedRepBench: Benchmarking Structured Understanding of Medical Report Images
Progressive and Localized Super-Resolution of 3D Objects via Localized Latent Voxel Diffusion
Minute4D: Training High-Fidelity 4D Gaussian Splatting in One Minute
Reconstructing Dense Depth of Dark Scenes with Sparse LiDAR, Noisy Events, and Blurry RGB
TaxoMIL: Taxonomy-Constrained Learning for Hierarchical Whole Slide Image Analysis
SurvMILKD: A Weakly Supervised Survival Analysis Framework for Multi-Teacher Knowledge Distillation using Pathology Foundation Models
Low-Level Dataset Distillation for Medical Image Enhancement
DualResPS: Dual-Resolution Photometric Stereo Using a Frame-Event Hybrid Camera
Fast and Accurate Image Restoration with Rank Enhanced Linear Attention
REVA-PO: Stabilizing Reinforcement Learning for Chest X-ray Report Generation
SHINE-PPG: Non-Lambertian Intrinsic Decomposition for Illumination-Robust rPPG
Physics-Guided Deep Learning for Linear Mueller Matrix Acquisition
Beyond the Embedding Bottleneck: Adaptive Retrieval-Augmented 3D CT Report Generation
Towards Reliable Multi-Label Classification via Conditional Dependency Modeling
Interpretation-Oriented Cloud Removal via Observation-Anchored Residual Flow with Geo-Contextual Alignment
Moonstone: A Multimodal Foundation Model and Benchmark for Lunar Remote Sensing
Structured SIR: Efficient and Expressive Importance-Weighted Inference for High-Dimensional Image Registration
Scale3D: Autoregressive Modeling for Large Outdoor Scene Generation
Vinci2: Providing Proactive Assistance in Continuous Egocentric Videos
Comprehensive language–image pre-training for 3D medical image understanding
iMED: A Multi-Endoscope Dataset for Surgical 3D Perception
GeoCFM: Positive-Only Conditional Flow Matching for Mineral Occurrence Sampling
ZipDepth: Bringing Lightweight Zero-Shot Monocular Depth Anywhere, on Any Device
OmniLife360: A Benchmark for 3D Reconstruction from In-the-Wild 360° Captures
MobileSAM2: Lightweight Segment Anything in Images and Videos via Hypergraphical Knowledge Distillation
3D Field of Junctions: A Noise-Robust, Training-Free Structural Prior for Volumetric Inverse Problems
NeuralDMD: Interpretable Untrained Neural Network for Imaging from Sparse and Noisy Observations
SEMIR: Topology-Preserving Graph Minors for Thin-Structure Segmentation
Pay Attention to Attention Distribution: A New Local Lipschitz Bound for Transformers
BiCE-HG: A Bi-Conditional Egocentric Hand Gesture Dataset for Intelligent Reality Systems
Learning from Adversity: Semantic-Aware Mask Refinement through Adversarial Perturbation
Adaptive Noise Covariance Scheduling under Riemannian Metrics for Diffusion Models
Escaping the Low-Frequency Bias: Adversarial Frequency Perturbation for Generalisable Gaze Estimation
Fully Rotation-Equivariant Spectral-Spatial Learning for Multispectral Object Detection
Don’t Teach Instability, Teach Robustness: Selective Sensitivity Gating for Adversarial Robust Distillation
One Scene, Two Depths: Probing Geometric Ambiguity in Monocular Foundation Models
TR-MoE: Temporal Reliability-Aware Mixture-of-Experts for Robust Tracking
REON-NVS: Real-Time Online Novel-View Synthesis from Sparse-View Videos
Generating Multi-view Adversarial Examples for Visual Geometry Grounded Transformer
WaterGen: Decoupling Scene and Medium in Underwater Image Generation
LoRC: Detecting AI-Generated Images via Low-Rank Collapse in the Semantic-Residual Space
Proteus: Model Leakage-Induced Adversarial Attack in Federated Learning
Beyond Filter Pruning: Top-K Spatial Selection for Efficient Neural Networks
MArFE: Multi-Contrast MRI Arbitrary Scale Super-Resolution with Fourier Enhancement
Explainability-aware Frustum Attack: Exposing Structural Vulnerabilities in LiDAR-Based 3D Object Detectors
Robustness Emerges Early in Training Dynamics, but Is Not Preserved
REVEAL: Reasoning-Enhanced Forensic Evidence Analysis for Explainable AI-Generated Image Detection
Implicit Neural Representation Facilitates Unified Universal Vision Encoding
SlowBA: An efficiency backdoor attack towards VLM-based GUI agents
Noise-Robust Face Recognition via Non-target Similarity Distribution Guided Sample Selection
UC-VLM: Consistency-Driven Learning for AI-Generated Image Detection with Vision-Language Large Models
SPICE: Simple Polysemantic feature Interpretation via Clustering-based Explanations
When Higher Order Hurts: Pre-Asymptotic Order Collapse in Generative ODE Sampling — A Theory of Discretization–Learning Interaction
BrainRiem: Riemannian Prototype Learning for Source-Free Cross-Site Brain Network Diagnosis
InSeg: Interactive Refinement via Intent Propagation for Point Cloud Semantic Segmentation
Bridging Theory and Practice in Source-Free Domain Adaptation via Adversarial Proxy Perturbation
SpaR3D-MoE: Adaptive 3D Spatial Reasoning from Sparse Views Meets Geometry-Inductive Mixture-of-Experts
Mapping the Concept Landscape: Structural Perception of Global Distributions for Transparent Data Pruning
Puppet-CNN: Continuous Parameter Dynamics for Input-Adaptive Convolutional Networks
DiffUE: Enhancing Utility-Unlearnability Trade-off of Unlearnable Examples via Diffusion Autoencoders
Learning with Bilevel-Minimax Optimization for Efficient and Reliable Transfer Attacks
Improving Adversarial Robustness by Mitigating Instability through Relearning
Prior-Conditioned Gaussian Discriminants for Generalizable AI-generated Image Detection
Dynamic-V2C: Editable and Continual Vision-to-Concept Bottleneck Models via Influence Functions
Evaluating the Interpretability of Sparse Autoencoders with Concept Annotations
Physically Grounded Dual-Opacity Gaussian Splatting for Joint RGB-TIR Reconstruction
Deep Noise Label Learning via Effective Rank Reduction
D-VLAM: Differential Vision and Language Mixing for Rehearsal Free Continual Learning
Adversarial Attack and Disturbance Detection by Hadamard-Coded Output Representations for Object Detection and Semantic Segmentation
Breaking High Confidence: Practical Face Impersonation under High-Security Thresholds
Multi-Anchor Distillation with Text-Guided Analytic Classifier for Continual Learning
FedMental: Topology-Aware Federated Prototype Learning for Polymorphic Multimodal Psychiatry
Guardrail-Agnostic Societal Bias Evaluation in Large Vision-Language Models
CS-TTA: Preserving Concept Sensitivity in Test-Time Adaptation
Efficient Quantization-Aware Adaptation for Visual Foundation Models
DnA: Denoising Attention for Visual Tasks
Boosting Text-Driven Video Segmentation via Geometry-Aware Distillation
Inductive Visual Logic for Few-Shot Out-Of-Distribution Adaptation in VLMs
Let ViT Speak: Generative Language-Image Pre-training
Quick ViTs: Speeding up Vision Transformers through Equivariance
Auto-Prompting: Layer-Specific Prompt Fusion Discovery via Differentiable Search
DeCoPatch: Revealing Causal Latent Subspaces in Vision-Language Models for GUI Grounding
PRISM3D: Probabilistic Refinement and Robust Initialization for Physically Consistent Scene Modeling under Extreme Motion Blur
Boosting Correspondence Learning with Structure-Aware Estimator
RBE-Flow:Recurrent Bayesian Estimation on Feature Manifolds for Cross-Modal Registration
Match-Any-Events: Zero-Shot Motion-Robust Feature Matching Across Wide Baselines for Event Cameras
XSurfer: Reconstructing surface meshes of cerebral and cerebellar cortex from diverse MRI data using untrained neural networks
Reflection-Aware Reasoning for Non-Line-of-Sight Pedestrian Localization
AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
Retrieving and Refining Winning Noise Tickets for Diffusion-Based Motion Generation
SGMatch: Semantic-Guided Non-Rigid Shape Matching with Flow Regularization
CtrlCoMo: Controllable Co-Speech Motion Generation with Gesture–Action Disentanglement
There and Back Again: A Flexible-Frame Transformer for Multi-Exposure Fusion
Optimizing Mesh Animation from Video via Shape Flow Guidance
LivingWorld: Interactive 4D World Generation with Environmental Dynamics
Sector-Level Cross-View Geo-Localization with Implicit Orientation via Azimuthal Scanning
Reasoning with Memory: A Temporal Granularity-Adaptive Framework for Training-Free Long Video Understanding
Category-Level Articulated Object Pose Estimation via Pose–Shape Hypothesis Generation and Verification
Every Dog Has Its Day, Probably: A Balanced Synthetic Benchmark and Probabilistic Modeling for 3D Dog Pose Estimation
PoseImageNet: Pose Estimation for Extensive Classes Based on Rich Structure Prototypes
Modeling and Compensating Phase Error in High-speed 3D Reconstruction
Flow4R: Unifying 4D Reconstruction and Tracking with Scene Flow
AnyMatch: Supercharging Universal Multi-Modal Image Matching with Large-Scale Single-View Images
SemLight: Distilled Semantic–Geometric Fusion for Efficient Local Feature Matching
Coordinate Singularities Break Conformal Coverage for Gaze and Head Pose
GrowFields: Compositional 4D Neural Fields for Topology-Changing Plant Growth
AirZoo: A Unified Large-Scale Dataset for Grounding Aerial Geometric 3D Vision
IoUCert: Robustness Verification for Anchor-based Object Detectors
Leaving the City: A Large-Scale Aerial Dataset for Cross-Season Localization in Unstructured Environments
Unveiling Transferability in Trajectory Prediction via Latent Scene Embeddings
PhyMAGIC: Physical Motion-Aware Generative Inference with Confidence-guided VLM
Trajectory-aware Cross-view Geo-Localization with Sequential Observations
StreetForward: Perceiving Dynamic Street with Feedforward Causal Dynamics
RegHead: Non-Humanoid Head Blendshapes via Feed-Forward Registration
UniDynamics: Event-RGB Fusion for Unified Future 4D Dynamic Scene Generation
UF0-6D: Unified Flow-based Zero-Shot 6D Object Pose Estimation without Refinement
Unified and Efficient Point-Line Local Features
Motion Style Slider: Endpoint-Supervised Continuous Style Control for Human Motion Diffusion
Degradation-Agnostic Clarity Learning for Unpaired Image Dehazing
ReflectCAP: Detailed Image Captioning with Reflective Memory
PolarAPP: Beyond Polarization Demosaicking for Polarimetric Applications
PrimitiveUDF: Primitive-Based Unsigned Distance Fields for Surface Reconstruction from Point Clouds
Embedding Rotation Invariance for Provable Multi-Oriented Scene Text Recognition
Discovering Geometric Biases in 3D Face Reconstruction: A Curvature-Aware Spectral Framework for Fairness Evaluation
TerrainGraphNet: Terrain-Constrained Graph Reasoning for Landslide Segmentation
Geometry-Preserving Image Generation for 6D Object Pose Estimation
Pix2NPHM: Learning to Regress NPHM Reconstructions From a Single Image
WiFi-JEPA: Self-supervised Learning for WiFi-CSI 3D Human Pose Estimation
PolyLayout: Multi-room Manhattan Layout Estimation
Visible Yet Unrecognizable: Frequency-Selective Facial Privacy via Attention
Learning Manifolds in High-D Point Embedding for Anisotropic Surface Approximation from Unstructured Point Clouds
CCFM: Collision-Constrained Flow Matching for Safety-Critical Scenario Generation
DisentangledTMR: Privacy-Preserving Skeleton Motion Retargeting via Factorized Transformers
PointSplat: Compact Gaussian Splatting via Human-Centric Prediction
DiGS-Avatar: Single-Image Animatable 3D Human Reconstruction via UV-Space Diffusion
Hi-DREAM: Brain Inspired Hierarchical Diffusion for fMRI-to-image Reconstruction via ROI Encoder And visual Mapping
Controlling Embedding Spaces with Text-Conditioned Transformations
DuoFlow: JVP-Free Finite-Difference Mean Flows for One-Step Image Generation
Affordance-Guided Diffusion Prior for 3D Hand Reconstruction
Dotting the Eye: An Intent-Driven Image Retouching Agent for Visual Focus Enhancement
Beyond the Black Box: Identifiable Interpretation and Control in Generative Models via Causal Minimality
Flow Straight to Reality: Perceptually Consistent Flow Matching for Efficient Image Restoration
Teaching an Agent to Sketch One Part at a Time
Volume Transformer: Revisiting Vanilla Transformers for 3D Scene Understanding
BRepFacetGen: Reverse Engineering B-Reps By Generative Face Segmentation
DOGE: Differentiable Bézier Graph Optimization for Road Network Extraction
JanusMesh: Fast and Zero-Shot 3D Visual Illusion Generation via Cross-Space Denoising
Geometric Foundation Model Distillation for Efficient Lunar 3D Reconstruction
Pathryoshka: Compressing Pathology Foundation Models via Multi-Teacher Knowledge Distillation with Nested Embeddings
REGLUE Your Latents with Global and Local Semantics for Entangled Diffusion
Parametric SDF for Dynamic Surface Reconstruction
STAT: Soft Tail-dropping for Adaptive Visual Tokenization
Unsupervised Pixel-Level Semantic Left-Right Understanding of In-the-Wild Images
PixelU: A U-Shaped Transformer for Efficient End-to-End Pixel Diffusion
TreeSRNF: Square-Root Normal Fields for Generative Modelling of the Geometric and Structural Variability in Tree-like 3D Objects
Temporally Stable Generative Illumination with a One-Step Diffusion Model
Learning to Tessellate: Point Cloud Generation via Recursive Spectral Partitioning
Learning Video Dynamics with Predictive Differentiable Rendering
DeCoFlow: Structural Decomposition of Normalizing Flows for Continual Anomaly Detection
CaRe: Critical Parameter Rectification for Efficient Visual Modeling
Hybrid-LUT: Channel-Aware Hybrid Lookup Table and Filtering for Efficient Image Restoration
WinTok: A Win-Win Hybrid Tokenizer via Decomposing Visual Understanding and Generation with Transferable Tokens
Pretrained Video Models as Differentiable Physics Simulators for Urban Wind Flows
WorldMesh: Generating Navigable Multi-Room 3D Scenes via Mesh-Conditioned Image Diffusion
Render-in-the-Loop: Vector Graphics Generation via Visual Self-Feedback
SAM2Matting: Generalized Image and Video Matting
Learning Geometry-Aware Embedding Fields for Intrinsic Riemannian Mappings
Thinking Ahead: Foresight Intelligence in MLLMs and World Model
Reflecting Process Expertise in Procedural Material Generation
AVQ-Attention: Adaptive Vector-Quantized Attention
DreamPartGen: Semantically Grounded Part-Level 3D Generation via Collaborative Latent Denoising
Language-Guided Transformer Tokenizer for Human Motion Generation
OmniX: From Unified Panoramic Generation and Perception To Graphics-Ready 3D Scenes
Reconstructing 3D Human-Object Interaction via a Unified Triplane Space
Geodesic Flow Matching on a Riemannian Degradation Manifold for Blind Image Restoration
Beyond Imitation: Learning Safe End-to-End Autonomous Driving from Hard Negatives
CoToGrasp: Contact-Topology-Conditioned Dexterous Grasp Synthesis via Canonical Workspace Learning
A Scalable Vector Graphics Latent Space
ProGVC: Progressive-based Generative Video Compression via Auto-Regressive Context Modeling
Monocular Avatar Reconstruction via Cascaded Diffusion Priors and UV-Space Differentiable Shading
Bridging the Geometry Mismatch: Frequency-Aware Anisotropic Serialization for Thin-Structure SSMs
ReFP-AD: Rectified Flow Preconditioning for Energy-Based Anomaly Detection
3DZip: Spatial-Aware Feature Diversity-Guided Token Compression for 3D Question Answering
Deformable and Multi-view Gradient-Aligned Physical Adversarial Camouflage
Through Van Gogh’s Eyes: Global Style Transfer with Diffusion Model
HQ-DM: Single Hadamard Transformation-Based Quantization-Aware Training for Low-Bit Diffusion Models
VNC: A Scale-Space Foundation for Learnable 3D Surface Evolution
Generalization and Memorization in Rectified Flow
LumiTokens: 3D Relighting via Token-Space Lighting Transformation
Q-BridgeNet: A Quantization Network for Cross-Lingual Sign Language Translation
Penetration-Free Compositional 3D Generation via Gaussian Surface Offset
Anatomy of a Lie: A Multi-Stage Diagnostic Framework for Tracing Hallucinations in Vision-Language Models
Diffusion Integrated Gradients: Controllable Path Generation for Flexible Feature Attribution
PRISM: Latent Composition Consistency for Single-Image Reflection Removal
DiffGI: Differentiable Geometry Images for High-Fidelity Thin-Shell 3D Generation
MagnetGS-Mesh: High-Quality Multi-Object Mesh Reconstruction via Adaptive Surface Optimization
WorldAgents: Can Foundation Image Models be Agents for 3D World Models?
K-Mask: Kinematic-Aware Masked Modeling for Controllable Text-to-Motion Synthesis
TiltDiff: Tilted Weight-Space Diffusion for Neural Network Generation
Why Linear Probing Works: Non-Vacuous Generalization Bounds via Effective Dimension
SWAN: World-Aware Adaptive Multimodal Networks for Runtime Variations
CMuon: Accelerating and Stabilizing Diffusion Transformer Training via Chunked Momentum Orthogonalization
NeuralGarSim: Geometry-agnostic Garment Simulation with Neural Fields
EoS-FM: Can an Ensemble of Specialist Models act as a Generalist Feature Extractor?
ExPLoRe: Expert Patch-Level Loss Routing for Multi-Objective Masked Image Modeling
InnoText: A Unified Model for Visual Text Generation and Editing
CoCo: Code as CoT for Text-to-Image Preview and Rare Concept Generation
PIPBench: A Profile-Inclusive Framework for Personalized Image Generation Evaluation
Layout-Conditioned Autoregressive Text-to-Image Generation via Structured Masking
NumColor: Precise Numeric Color Control in Text-to-Image Generation
Vision Bridge Transformer at Scale
Teaching Vision-Language-Action Models What to See and Where to Look
Learn Once, Edit Anywhere: Visual Direction Transfer for Diffusion Models
FLM-Occ: Feed-forward Likelihood Maximization for Efficient Indoor Occupancy Prediction
VGEdit: Unlocking Video Generation Priors for Reasoning-Informed Image Editing
Anchoring on Reality: Breaking the Pseudo-Target Ceiling in Makeup Transfer
Aligning Human Sense: Calibrated Distributional Reward Learning for Video Generation
Reliable Reasoning in SVG-LLMs via Multi-Task Multi-Reward Reinforcement Learning
MANGO: Unleashing Image Generation Capability of Unified Multimodal Models
Anti-Prompt: Image Protection against Text-Guided Image-to-Video Generation
LinCa: Accelerating Diffusion Models via Learnable Decomposed Feature Caching
UEval: A Benchmark for Unified Multimodal Generation
LISA: Locality-Informed Speculative Decoding for Accelerating Autoregressive Image Generation
High-Throughput Event-Based Feature Detection and Tracking on an Embedded CPU
RePlan: Reasoning-Guided Region Planning for Complex Instruction-Based Image Editing
UNet-Twice: A Simple Structured Reference-based Inpainting Framework
Scientific Image Synthesis: Benchmarking, Methodologies, and Downstream Utility
SAMA: Factorized Semantic Anchoring and Motion Alignment for Instruction-Guided Video
LensStyle: Learning the Optical Aesthetics for Controllable Stylized Lens Effect Rendering
TRaM-VSR: Importance-Aware Token Routing and Merging for One-Step Diffusion Video Super-Resolution
In-context Region-based Drag: Drag Any Region to Any Shape
FlowInOne: Unifying Multimodal Generation as Image-in, Image-out Flow Matching
AR-CoPO: Align Autoregressive Video Generation with Contrastive Policy Optimization
PriorEye: Geospatial Visual Priors for End-to-End Autonomous Driving
Multi-dimensional Preference Alignment by Conditioning Reward Itself
Follow-Your-Mind: Towards Inversion-Free Brain-Driven Visual Context Synthesis and Editing
Histocomponent-driven Universal Model for Virtual Immunohistochemistry Multiplex Staining via Joint Manifold Evolution
DRPO: Disentangling Demographic Bias from Rewards for Fair Diffusion Alignment
Beyond Fixed Luminance: Towards Panchromatic and Orthochromatic Image Colorization
EruDiff: Refactoring Knowledge in Diffusion Models for Advanced Text-to-Image Synthesis
VTEdit-Bench: A Comprehensive Benchmark for Multi-Reference Image Editing Models in Virtual Try-On
Training-Free Multi-Concept Image Editing
DARE to Mitigate Hallucination: Dual-path Auto-Regressive-aware Editing
Q-REAL: Towards Naturalness and Distortion Evaluation for AI-Generated Content
RankT2I: A Submodular Framework for Discovering Interpretable and Diverse Semantics in Text-to-Image Models
Dress-ED: Instruction-Guided Editing for Virtual Try-On and Try-Off
From Open Loop to Closed Loop: A Test-Time Iterative Optimization Framework for Reference-Consistent Image Generation
ScrollScape: Unlocking 32K Image Generation With Video Diffusion Priors
Attribute Token Arithmetic: Disentangled and Continuous Semantic Control for Visual Autoregressive Models
GenAgent: Scaling Text-to-Image Generation via Agentic Multimodal Reasoning
IRG-MotionLLM: Interleaving Motion Generation, Assessment and Refinement for Text-to-Motion Generation
Robustness Meets Uncertainty: Evidential Adversarial Training for Robust Selective Classification
Revisiting Autoregressive Models for Generative Image Classification
A²-Edit: Precise Reference-Guided Image Editing of Arbitrary Objects and Ambiguous Masks
Flash-BoN: Instant Drafts for Inference-Time Scaling in Diffusion Models
RubricRL: Simple Generalizable Rewards for Text-to-Image Generation
ELDiff: When Evidential Learning Meets Text-to-Image Diffusion
Rethinking Garment Conditioning in Diffusion-based Virtual Try-On: Decouple, Don't Denoise
Fast and Scalable LiDAR Data Generation for Autonomous Driving Simulation without Raycasting
Layering Virtual Try-On
PosterCopilot: Toward Layout Reasoning and Controllable Editing for Professional Graphic Design
Setting the Stage: Text-Driven Scene-Consistent Image Generation
PPTArena: A Benchmark for PowerPoint Editing
In-Context Sync-LoRA for Portrait Video Editing
RADIANCE: Relative Adaptive Denoising with IP-Adapter for Novel Concept Enhancement
UniREditBench: A Unified Reasoning-based Image Editing Benchmark
SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models
Policy-Based Tuning of Autoregressive Image Models with Instance- and Distribution-Level Rewards
RiO-DETR: DETR for Real-time Oriented Object Detection
Spanning the Visual Analogy Space with a Weight Basis of LoRAs
Stream-DiffVSR: Low-Latency Streamable Video Super-Resolution via Auto-Regressive Diffusion
Spanning Tree Autoregressive Visual Generation
HRDiT: Training-Free High-Resolution Image Generation with Off-the-Shelf Diffusion Transformer Models
POET: Preference Optimization for Enhanced Text-to-Image Generation
Discrete Noise Inversion for Next-scale Autoregressive Text-based Image Editing
TOPA: Mitigating Concept Dominance in Diffusion Personalization via Target-Oriented Perturbation Augmentation
Reward Lightning: Fast Video Generation via Homologous Preference Distillation
Reinforcement Learning for Multimodal Diffusion Language Models via Bidimensional Trajectory and Thought Optimization
Direct Diffusion Score Preference Optimization via Stepwise Contrastive Policy-Pair Supervision
DreamLite: A Lightweight On-Device Unified Model for Image Generation and Editing
OmniMamba: Efficient and Unified Multimodal Understanding and Generation via State Space Models
Personalized Reward Modeling for Text-to-Image Generation
Distribution Matching Distillation Meets Reinforcement Learning
Tiled Prompts: Overcoming Prompt Misguidance in Image and Video Super-Resolution
Unmasking-Time Visual Calibration for Hallucination Mitigation in Multimodal Discrete Diffusion Language Models
GR-GRPO: Graph-Diffused Credit for Autoregressive Image RL Alignment
LayerVerse: Finding the Sweet Spot for KV-Injection in Training-Free Image Editing
TerraDiT-Ω: Unified Spatial Control for Satellite Image Synthesis with Any Geospatial Primitive
Achieving Subcategorical Erasure in Text-to-Image Models
FairSteer: Cross-Attention Steering Towards a Fairer Text-Guided Image Generation
Posterior Augmented Flow Matching
Unsafe by Reciprocity: How Generation–Understanding Coupling Undermines Safety in Unified Multimodal Models
WearWow: Native 2K Multi-Garment Virtual Try-On via Adaptive Token Packing and Preference Alignment
Target-aware Image Editing via Cycle-consistent Constraints
FineEdit: Fine-Grained Image Edit with Bounding Box Guidance
SpecV: Specification Verification for Robust Unified Multimodal Evaluation
CGCE: Classifier-Guided Concept Erasure in Generative Models
Beyond Absolute Scores: Relative Edit-induced Difference for Generalizable Image Aesthetic Assessment
FDM-MFVT: Few-step Sampling Diffusion Model for Mask-Free Virtual Try-On
DC-Gen: Post-Training Diffusion Acceleration with Deeply Compressed Latent Space
Illuminating Unified Multimodal Model for Free-form Interleaved Text-Image Generation
RefReward-SR: LR-Conditioned Reward Modeling for Preference-Aligned Super-Resolution
Omni-RRM: Advancing Omni Reward Modeling via Automatic Rubric-Grounded Preference Synthesis
Reinforcing Video Reasoning with Focused Thinking
Obliviate: Erasing Concepts from Autoregressive Image Generation Models
Not All Prediction Targets Keep Training-Free Diffusion Guidance on the Manifold
Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning
LightFusion: A Light-weighted, Double Fusion Framework for Unified Multimodal Understanding and Generation
OTCache: Optimal Transport for Geometry-Aware Caching in Diffusion Models
Diffusion-SDPO: Safeguarded Direct Preference Optimization for Diffusion Models
PhyEditBench: A Real-World Multi-Stage Benchmark for Physics-Aware Image Editing
MagicMakeup: A Region-Controllable Diffusion Transformer for High-Fidelity Makeup-Transfer
Wavelet-Guided Semantic Signal Compensation for Inversion-Free Image Editing
Rethinking Robust Adversarial Concept Erasure in Diffusion Models
StereoEdit: A Diffusion-Based Framework for Stereo-Consistent Image Editing
Unlocking Complex Image Editing via Natively Interleaved Visual Textual CoT with Deep Confidence Reasoning
OSVE: One Step Video Editing with One Step Diffusion Models
InstaEdit: Instant Image Editing via Optimized Noise Prediction
TexTailor: Inference-Time Textual Guidance Tailoring for Multimodal Diffusion Transformers
Coding with Eyes: Visual Feedback Unlocks Reliable GUI Code Generating and Debugging
PhysEdit: Physically Consistent Image Editing via Causal Enforcement
MVI2V: Human Centric Image to Video Generation with Multiview Consistent Appearance
Zero-Shot Image Personalization from Personas
The Path to Reconciling Quality and Safety Alignment in Text-to-Image Generation
GIDE: Unlocking Diffusion LLMs for Precise Training-Free Image Editing
AVSR-Diff: Scale-Agnostic Diffusion Priors for Temporally Consistent Arbitrary-Scale Video Super-Resolution
Optimization-Guided Diffusion for Interactive Scene Generation
OSOR: One-Step Diffusion Inpainting for Effect-Aware Object Removal
Reference-Free Quality Assessment for Virtual Try-On via Human Feedback
CellFluxRL: Biologically-Constrained Virtual Cell Modeling via Reinforcement Learning
SCALE: Semantic-Calibrated Guidance Enhancement for Prompt-Faithful Diffusion
Seeing to Ground: Visual Attention for Hallucination-Resilient MDLLMs
Intermediate Text Representation Guided Text-to-Image Generation for Enhancing One-and-Only Alignment
Push–Pull Attentional Anchoring for Diffusion Concept Erasure
SPQR: A Multi-Dimensional Benchmark for Safety Alignment under Benign Model Adaptation
MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE
SparkVSR: Interactive Video Super-Resolution via Sparse Keyframe Propagation
Multi-History-Step SDE Inversion for Image Editing with Superior Regional Awareness
H-Adapter: Pose-Robust Hairstyle Transfer via Attention-Derived, Source-Aligned Hair Masks
MirrorPPR: Exemplar-Based Portrait Photo Retouching
Joint Alignment and Distillation for Video Generation via Sample-Guided Distribution Matching
Syn-GRPO: Self-Evolving Data Synthesis for MLLM Perception Reasoning
When Rubrics Fail: Error Enumeration as Reward for Reference-Free RL Post-Training
EraseSAE: Surgical Concept Erasure in Text-to-Video Diffusion Models via Sparse Autoencoders
Contextual Image Attack: How Visual Context Exposes Multimodal Safety Vulnerabilities
Continuous Speculative Decoding for Autoregressive Image Generation
SciIR: A Large-scale Training Dataset and Benchmark for Scientific Image Reasoning Generation
Editing Everything Everywhere All at Once
Learning Consistency in Reward Modeling for Multi-Modal Reasoning
OpenSubject: Leveraging Video-Derived Identity and Diversity Priors for Subject-driven Image Generation and Manipulation
SR-Edit: Region-Aware Image Editing via Self-Refinement
Wan-R1: Verifiable-Reinforcement Learning for Generalizable Video Reasoning
Semantically Aligned Gradient-Driven Context-Preserving Image Editing
When Cars Have Stereotypes: Auditing Demographic Bias in Objects from Text-to-Image Models
FED-Bench: A Cross-Granular Benchmark for Disentangled Evaluation of Facial Expression Editing
VLTR: Vision-Language Tool Reasoning for Instruction-Guided Image Editing
Drift-AR: Single-Step Visual Autoregressive Generation via Anti-Symmetric Drifting
To Erase, or Not to Erase: Robust Training-Free Concept Erasure with Preservation aware Adaptive Ranked Subspace Expansion
The Map Is Not the Territory: Embedding-Coverage Blacklists for Safe Diffusion Steering
M4-SAR: A Multi-Resolution, Multi-Polarization, Multi-Scene, Multi-Source Dataset and Benchmark for optical-SAR Object Detection
Decoding Children’s Gait Behavior
InclusiveHuman-10K: Towards Inclusive Human Parsing Beyond the Intact-Limb Assumption
EDM:Event-guided Diffusion Model for Video Shadow Detection in Complex Dynamic Scenes
Streaming Dense Voxel Representations for 3D Occupancy Prediction
Learning Probabilistic Embeddings for Unsupervised Action Segmentation
Region-Aware Multimodal Interleaving for Animal Re-Identification
SGQA: Semantic-Geometric Quality Alignment for Training-Free Few-Shot Instance Segmentation
Fast Dynamic Prototypes for Unsupervised Anomaly Detection and Localization
Free‑CD: Probabilistically Decoupled Training-Free Open-Vocabulary Change Detection with Resolution-Invariant Feature Inversion
Silhouette-based Gait Foundation Model
SVCBench: A Streaming Video Counting Benchmark for Spatial-Temporal State Maintenance
STEP: Score-Based Temporal Energy for Human Pose Video Anomaly Detection
Rethinking Continual Anomaly Detection on the Edge: Benchmarking Under Realistic Industrial Conditions
FoundYou: A Unified Model for Personalized Segmentation and Retrieval
Back-Tracking from Clarity: Self-Learning to See Text from Afar
RePL: Pseudo-label Refinement for Semi-supervised LiDAR Semantic Segmentation
One Slide, Many Views: Unifying Complementary Foundation Model Perspectives for WSI Analysis
MiNQVIS: Mitigating Noisy Queries for Robust Online Video Instance Segmentation
LlamaSeg: Image Segmentation via Autoregressive Mask Generation
PS-MOT: Cultivating Instance Awareness from Point Seeds for Multi-Object Tracking
CMCC-ReID: Cross-Modality Clothing-Change Person Re-Identification
DeMuS: Learning Decoupled Matching and Scoring for Batch Zero-Shot Industrial Anomaly Detection
Uncertainty-aware tree height change regression
Hierarchical Hyperbolic Representation Learning for Aerial-Ground Person Re-Identification
Beyond Attention: Convolutional Global Context for Remote Sensing Change Detection
LeAD-M3D: Leveraging Asymmetric Distillation for Real-Time Monocular 3D Detection
ELHINN: Unifying Dense Crowd Simulation Across Scales via Eulerian–Lagrangian Hydrodynamics
CountEx: Fine-Grained Counting via Exemplars and Exclusion
GroundingAnomaly: Spatially-Grounded Diffusion for Few-Shot Anomaly Synthesis
InstaPano: Zero-shot Instance Layout Controlled Panorama Generation Via Global Attention Fusion
Where and What: Long-Term Object Tracking in Egocentric Videos
Efficient RGB-T Object Detection via Sparse Cross-Modality Fusion
ARGOS: Who, Where, and When in Agentic Multi-Camera Person Search
SCORE: SubDistribution-aware Collaborative Knowledge Reinforcing for Cloth-Hybrid Lifelong Person Re-Identification
SegVGGT: Joint 3D Reconstruction and Instance Segmentation from Multi-View Images
Solving Semi-Supervised Few-Shot Learning from an Auto-Annotation Perspective
Generating a Paracosm for Training-Free Zero-Shot Composed Image Retrieval
Divide and Align: Disentangled Vision-Language Learning for Text-Based Person Anomaly Search
IREU: Identity-Related Encoder-Only Unlearning for Customized Portrait Generation
EM3M: An Electron Micrograph Dataset for Microstructural Segmentation and Generation
WebEyeTrack: Scalable Eye-Tracking for the Browser via On-Device Few-Shot Personalization
EVEE: Event-Based Online Adaptation for Matching on Unknown Targets
TMI: Text-to-Image Meets Image-to-Image for Complementary Data Synthesis to Boost Long-Tailed Instance Segmentation
Seek to Segment: Active Perception for Panoramic Referring Segmentation
FusionTrack: Collaborative Multi-Object Tracking with Arbitrary Multi-UAVs
Environmental Change Detection for Real-World Change Analysis
SplitHDR: Saturation-Aware HDR Recovery and Denoising for Real-Time Detection
DeltaDeno: Zero-Shot Anomaly Generation via Delta-Denoising Attribution
History-Aware Transformation of ReID Features for Multiple Object Tracking
Beyond Aesthetics: Quantifying Information Loss in Turbid Scenes
Unsupervised Source-Free Ranking of Biomedical Segmentation Models Under Distribution Shift
Beyond Common Sense: Grounding Logical Anomaly Detection in Inspection Criteria
P3-SAM: Native 3D Part Segmentation
Low-latency Event-based Object Detection with Spatially-Sparse Linear Attention
Per‑Object IoU Forecasting for Deadline‑Aware Real‑Time Embedded Detection Control
Compositional Non-Face Re-Identification Pressure under Cumulative Vision Releases
StAR: Segment Anything Reasoner
ZMIS-SAM: Segment Anything Model Enhanced With Wavelet Transform For Zooplankton Microscopy Image Instance Segmentation
TiCRL: Textual Image Classification with Reinforcement Learning-Based Curriculum Learning
Slim-DETR: Real-Time Tiny Object Detection with Efficient Interaction and Gaussian Query
O-VAD: Industrial Video Anomaly Detection through Object-Centric Tracking and Reasoning
PA-VAD: Diffusion-Based Pseudo-Only Video Anomaly Detection via Domain-Aligned Memory Updates
DDStereo: Efficient Dual Decoder Transformers for Stereo 3D Road Anomaly Detection
Align and Segment: Unsupervised Learning for Building Segmentation From Misaligned Labels
ODONet: Online Dynamic Offset Network for Visual Object Tracking
BrainFIBRE: A Foundation Model via Information Decomposition for Brain Microstructure
MCVL: Multi-Space Cross-View Learning for Aerial-Ground Person Re-Identification
HieDG: A Hierarchical Discrete Geometry-Guided Framework for Multi-Animal Tracking
UnderOneFacade: Worldwide Facade Semantic Segmentation Benchmark Dataset
Event Stream-based Sign Language Translation: A High-Definition Benchmark Dataset and A Novel Baseline
MATCH: Flow Matching for Multi-View Anomaly Detection
Evidence Triangulation for Multimodal Fact-Checking in the Wild
Constrained Rotation Optimization: Revisiting Crop-Based Gaze Estimation
HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark
Proximity-CLIP: Text-Guided Semantic Proximity Learning for Zero-Shot Anomaly Detection
Uncertainty-Weighted Fusion of Image and Synthetic Event for Video Anomaly Detection
Real-Time Source-Free Object Detection
DETRAM: End-to-end DEtection, Tracking and Recovery of HumAn Meshes
Harnessing SSL for Segmentation in 3D Microscopy with Noisy Labels and Hard Patches
Anomaly Factory 3D: A Modular Framework for Diverse Pseudo-Anomaly Synthesis in Unsupervised 3D Anomaly Detection
Seeing as Humans Do: Learning from Motion to Segment Anything Without Supervision
SAGE: A Synchronized Action and Gaze Estimation Framework for Comprehensive Human Behavior Analysis
FuDU: A Fuzzy Dual-dimension Uncertainty Framework for Streaming Active Learning in Industrial Defect Detection
InstanceControl: Controllable Complex Image Generation without Instance Labeling
ANFI: Rethinking Neighbor Feature Interaction in Person Re-ID
ArcAD: Anomaly-Rectified Calibration for Cold-Start Supervised Anomaly Detection
Instance Segmentation as Tracking: A New Paradigm for Multi-Small-Object Tracking with Event Cameras
Mode-Conditioned Residual Calibration for Multi-Object Tracking
Domain Adaptive Object Detection via Dual-Stream Bilevel-Cycle Optimization
HLRAD: High-dimensional Latent Representation for Unified Anomaly Detection
Towards Video Anomaly Detection from Event Streams: A Baseline and Benchmark Datasets
WAPR: A Foundation Model for Wide-Angle Refinement in Unseen Object Pose Estimation
Local Spacing-Aware Hungarian Matching for Stable Point-Supervised Crowd Counting
SARIF: Segment Anything for Robust Image Forensics
A Comprehensive Analysis about Unsupervised Outlier Detection for Images
On-Orbit Real-Time Wildfire Detection Under On-Board Constraints
A Simple Baseline with Placement Prior for Point-Supervised Oriented Object Detection
RT-DocLayout: Real-Time End-to-End Document Layout Analysis with Reading Order in the Wild
ISAC: Training-Free Instance-to-Semantic Attention Control for Multi-Instance Generation
RT-RMOT: A Dataset and Framework for RGB-Thermal Referring Multi-Object Tracking
3D-LENS: A 3D Lifting-based Elevated Novel-view Synthesis method for Single-View Aerial-Ground Re-Identification
Robust Zero-shot Anomaly Detection under Limited Auxiliary Anomaly Priors
Reliability-Aware 3D Geometric Injection for Universal Person Re-identification
Denoising-Enhanced Coarse-to-Fine Infrared Small Target Detection with Attention Prior-Guided Knowledge Distillation
COLA: Continual Orthogonal Low-Rank Adaptation for Class-Incremental Learning
BAAF: Universal Transformation of One-Class Classifiers for Unsupervised Image Anomaly Detection
MEVL-STP: Multi-Encoder and Vision Language Model for Arbitrarily Shaped Scene Text Spotting
ModTrack: Sensor-Agnostic Multi-View Tracking via Identity-Informed PHD Filtering with Covariance Propagation
PhenoLeaf-TS: A Time-Series Benchmark for Leaf Instance Segmentation, Tracking, and Growth Stage Classification
Diagnosing Aerial-View Object Detectors with Foundational Image Generative Models
CRD-Net: Frequency-Adaptive Feature Injection and Change Decoupling for Building Damage Assessment
CGCC: Towards Generalizable Clothes-Changing Person Re-Identification
Segmenting, Fast and Slow: Real-Time Open-Vocabulary Video Instance Segmentation with Dual-Path Processing
Cross-Species Animal Re-Identification with Semantic Consistency Learning
Following the Flow: Advection-Consistent Modeling for Event-based Small Object Detection
Bounding-Box Trajectories Matter for Video Anomaly Detection
EVKit: An Open-source Flexible Toolkit for Efficient Event Camera Data Storage and Loading
ET-SAM: Efficient Point Prompt Prediction in SAM for Unified Scene Text Detection and Layout Analysis
Event-based Gaze Control Systems for Real-time Spin Estimation in Professional Ball Games
VIGA: View-Conditioned and Identity-Guided Adaptation for Aerial-Ground Person Re-Identification
Point2Pose: Occlusion-Recovering 6D Pose Tracking and 3D Reconstruction for Multiple Unknown Objects Via 2D Point Trackers
ProtoFair: Fair Self-Supervised Contrastive Learning via Pseudo-Counterfactual Pairs
Segmentation-Guided Homography Estimation for Long-Term Planar Tracking
Vitality-Aware Compression for Efficient Image-to-Shape Diffusion Transformers
SK-Adapter: Skeleton-Based Structural Control for Native 3D Generation
ARVAR: Accelerating Visual Autoregressive Model via Attention Retrospect
ViBe: Ultra-High-Resolution Video Synthesis Born from Pure Images
LatSearch: Latent Reward-Guided Search for Faster Inference-Time Scaling in Video Diffusion
Generative Refinement Network for Visual Synthesis
When Distillation Breaks Motion Control: Restoring Generative Trajectories for Fast Video Generators
HNDiff: Haze-Noise Diffusion for Image Dehazing
QualiTeacher: Quality-Conditioned Pseudo-Labeling for Real-World Image Restoration
From smooth to sharp: Frequency-Decoupled Latent Optimization for Realistic Image Generation
Learning to Corrupt for Better Restoration
Physics Meets Perception: A Reinforcement Learning Framework for Unpaired Real-World Image Dehazing
RefAlign: Representation Alignment for Reference-to-Video Generation
Glance: Accelerating Diffusion Models with 1 Sample
Analyzing and Improving Training-Free Fast Sampling of Text-to-Image Diffusion Models
Towards Scalable Pre-training of Visual Tokenizers for Generation
Adversarial Score Distillation for Stable One-Step Diffusion in Real-World Image Super-Resolution
SD3.5-Flash: Distribution-Guided Distillation of Generative Flows
From Noise to Events: Conditional Diffusion for Event Data Augmentation
Region-Aware Test-Time Scaling for Compositional Image Generation
Decoupling Complexity from Scale in Latent Diffusion Model
Beyond the Boundary: RL-Driven Solution Space Exploration for Blind Face Restoration
EMAG: Self-Rectifying Diffusion Sampling with Exponential Moving Average Guidance
QWERTY: Training-Free Motion Control via Query-Warped Video Diffusion Transformers
Histogram-constrained Image Generation
RTE-FM-Dehazer: Radiative Transfer Equation Inspired Flow Matching for Real-World Image Dehazing
Parsimonious Flow Matching for Efficient Image Generation
AlignMorph: Tuning-Free Diffusion Image Morphing via Explicit Semantic Transport
DiverseVAR: Balancing Diversity and Quality of Next-Scale Visual Autoregressive Models
SkelEM: Explicit Decoupling of Topology and Details for Self-supervised Axial Super-Resolution in Volume Microscopy
Recolour What Matters: Region-Aware Colour Editing via Token-Level Diffusion
InverseCrafter: Efficient Video ReCapture as a Latent Domain Inverse Problem
Fill2SR: Repurposing Inpainting Diffusion Transformers for Real-World Super-Resolution
Reconstruction-Anchored Diffusion Model for Text-to-Motion Generation
Moebius: 0.2B Lightweight Image Inpainting Framework with 10B-Level Performance
CMDer: Controllable Mode Decomposition-Based Single Motion Synthesis with Diffusion
SSDD: Single-Step Diffusion Decoder for Efficient Image Tokenization
Learning on the Manifold: Unlocking Standard Diffusion Transformers with Representation Encoders
Solving Diffusion Inverse Problems with Restart Posterior Sampling
Contrastive Conditional–Unconditional Alignment for Long-tailed Diffusion Model
GLARE: Towards Generalizable Detection of Latent Diffusion Images with Global-Local Reconstruction Error
ColorFM: An Optimization-to-Learning Framework for Color Transfer via Flow Matching
Be Tangential to Manifold: Discovering Riemannian Metric for Diffusion Models
Straight-Path Flow Matching for Incomplete Multi-View Clustering
ART-VSR: Adaptive Rectified Trajectories for One-Step Video Super-Resolution
Zero-Shot Inference-Time Rectification for Real-World Arbitrary-Scale Super-Resolution
GRAN-TED: Generating Robust, Aligned, and Nuanced Text Embedding for Diffusion Models
OARS: Process-Aware Online Alignment for Generative Real-World Image Super-Resolution
G3AFT: Glance Guided Gradient Aligned Fine-Tuning for Visual Autoregressive Models
UniCSG: Unified High-Fidelity content-constrained style-driven generation via Staged Semantic and Frequency Disentanglement
Hi-DiT: Hybrid Latent-Pixel Diffusion Transformer for Image Generation
Cross-Space Distillation: Teaching One-Step Students with Modern Diffusion Teachers
MTVDiff: Multimodal Conditional Latent Diffusion for Enhanced Thermal-to-Visible Face Translation
SAND: Stage-Aware Noise Decomposition for Training-Free Diffusion Guidance
AnyStyle: A Single LoRA is Sufficient for Image-Guided Style Transfer
Rethinking Real-World MRI Denoising: Learning from Physical Noise
FUMO: Prior-Modulated Diffusion for Single Image Reflection Removal
MaterialFlow: Attribute-Disentangled Material Transfer via Trajectory-Aware Velocity Modulation
Diffusion Image Generation with Explicitly Modeling of Data Manifold Geometry
Linear Fusion MultiDiffusion for Fast Training-Free Spherical Panorama Generation
Introspective Attention Modulation for Safe Text-to-Image Generation
Trust-Region Noise Search for Black-Box Alignment of Diffusion and Flow Models
Continuous Adversarial Flow Models
ReAL: Reference-to-Image (R2I) Aware Latent Diffusion for Image Super-Resolution
Fair and Faithful: A Diffusion-Enhanced Dataset and Hybrid State-Space Mamba for Face Super-Resolution
High-Resolution Artwork Outpainting with Global Blueprint Guidance and Layout Control
ConceptWeaver: Weaving Disentangled Concepts with Flow
Beyond Linear Shortcuts: Rectifying Diffusion Preference Optimization with Intrinsic Generative Geometry
D2PO: Optimizing Diffusion Samplers via Dynamic Preference
AC3S: Adaptive Conditioning for 3D-Aware Synthetic Data Generation
Spectral Prior for Reducing Exposure Bias in Diffusion Models
Y-diff: Structure-Texture Decoupled Diffusion Distillation for H&E-to-pCLE Translation
DiffRGD: An Inference-Time Diffusion Guidance Through Riemannian Gradient Descent
Dual-End Consistency Model
GeoFlow: Efficient Driving Video Generation via Geometry-Aligned Priors
Controllable Generative Reference for Stereo Image Compression via Reliability-Aware Gating
D²R²OSR: Degradation-Disentangled Representation for Real-World Omnidirectional Image Super-Resolution
Towards Consistent and Efficient Dataset Distillation via Diffusion-Driven Selection
Cross-Resolution Distribution Matching for Diffusion Distillation
RhymeFlow: Training Free Acceleration for Video Generation with Asynchronous Denoising Flow Scheduling
FMA-Net++: Motion- and Exposure-Aware Joint Video Super-Resolution and Deblurring
Anchoring and Steering Diffusion: Enhancing the Faithfulness of Text-to-Image Generation at Inference Time
PixGS: Pixel-Space Diffusion for Direct 3D Gaussian Splat Generation
Exposure Bias Can Alleviate Itself via Directional and Frequency Rectification in Flow Matching
BiSLW: Bi-Spectral Latent Watermarking for Generative Diffusion Models
EFlow: Fast Few-Step Video Generator Training from Scratch via Efficient Solution Flow
Co-evolving Representations in Joint Image-Feature Diffusion
Extreme Face Super-Resolution through Identity Fitting and Decoupling
Momentum Guidance: Plug-and-Play Guidance for Flow Models
Which Layer Causes Distribution Deviation? Entropy-Guided Adaptive Pruning for Diffusion and Flow Models
Curvature-Adaptive Consistency Flow Matching: Autonomous Trajectory Optimization via Reinforcement Learning
Tuning Real-World Image Restoration at Inference: A Test-Time Scaling Paradigm for Flow Matching Models
ECTraj: Enhanced Consistency Training for Multi-Agent Trajectory Prediction
Short-to-Long Functional Connectivity Transfer via Structure-Aware Latent Diffusion
RefDiT: Local Attribute Guidance in Reference-Based Image Generation
LUCE: Constrained Curve-Domain Guidance for Training-Free Low-Light Enhancement with Hue-Preserving Decoupling
Improving Image-to-Image Translation via a Rectified Flow Reformulation
DTI: Dynamic Trajectory Initialization for Generative Face Video Super-Resolution
ELT: Elastic Looped Transformers for Visual Generation
AdaBridge-SR: Adaptive Bridge Matching for Real-World Image Super-Resolution
Rethinking Token Reduction for Diffusion Models via Output-Similarity-Awareness
Trajectory Forcing: Structure-First Generation with Controllable Semantic Trajectories
Early Estimation of Language to Latent Alignment in Diffusion Models
FMS2: Unified Flow Matching for Segmentation and Synthesis of Thin Structures
DICT: Data Injection and Contrastive Trajectory Refinement for Conditional Image Generation with Diffusion Models
LUA: Latent Upscaling Adapter for Diffusion-Based Image Synthesis
Learning to Balance: Decoupled Siamese Diffusion Transformer for Reference-Based Remote Sensing Image Super-Resolution
Steering Diffusion Models via Class-Contrastive Influence for Few-Shot Classification
Accelerating Diffusion Transformers with Gaussian Process Rectified Feature Cache
DAPS++: Rethinking Diffusion Inverse Problems with Decoupled Posterior Annealing
One-Step Flow Policy: Self-Distillation for Fast Visuomotor Policies
Generative Manifold Distillation: Aligning Restoration Trajectories with the Natural Image Prior
Thermo-JEPA: Learning a Geometry-Grounded Thermal World Model via Cross-Modal Privileged Masking
Cross-token Guidance Transformer for Weakly Supervised Object Localization
Adaptive Spectrum-Aware Feature Disentangled Network for Small Object Detection
ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding
OpenPanoD: Aligning Multimodal Prompts and Spherical Representations for Open-Vocabulary Panoramic Detection
Multi-modality Image Fusion under Adverse Weather: Mask-Guided Feature Restoration and Interaction
CoT-PL: Chain-of-Thought Pseudo-Labeling for Open-Vocabulary Object Detection
RT-SDGOD: Real-Time Single-Domain Generalized Object Detection
A Dual-space Patch-driven Complementary Learning Framework for Semi-supervised Multi-organ Segmentation
Selective Synergistic Learning for Video Object-Centric Learning
Segmenting Visuals With Querying Words: Language Anchors For Semi-Supervised Image Segmentation
Disentangling and Reusing Interaction Cues for Zero-Shot HOI Detection
BEVOpen3D: Towards Open-World 3D Object Detection in Bird's-Eye-View
SGP2: Coarse-to-Fine Controllable Multimodal Remote Sensing Image Generation
CloSeR: Unified Relational Distillation from Closed-Set Teachers for Category Discovery
UniMotion: A Unified Framework for Motion-Text-Vision Understanding and Generation
HilDA: Hierarchical Distillation with Diffusion for Advancing Self-Supervised LiDAR Pre-training
ReMoMask: Retrieval-Augmented Masked Motion Generation
Decoding Multimodal Causality: End-to-End Multimodal Mediation Pathways Inference
T-REN: Learning Text-Aligned Region Tokens Improves Dense Vision-Language Alignment and Scalability
VarProtoAD: Variational Prototype-Conditioned Prompting for Zero-Shot Anomaly Detection
What Images Cannot Say: Language-Guided Olfactory Representation Learning
Distribution-Aware Feature Selection for Post-hoc Out-of-Distribution Detection
Incentive Noise and Structural Prior Infusion for Multi-Modal Object Re-Identification
DAP: Doppler-aware Point Network for Heterogeneous mmWave Action Recognition
FSD-Net: Foundation-Guided Spatiotemporal Distillation for Video Polyp Segmentation
CMDS-AD: Cross-Modal Dual-Stream Decoupling for Few-Shot Anomaly Detection
PhysFlowNet: Learning Canonical Latent Manifolds via Spatio-Spectral Physics Priors for Underwater Object Detection
DA-F2F: Domain-Adaptive Object Detection with Feature-to-Feature Modulation and Alignment
Intra-Class Consistency Guided Class-Agnostic Event Segmentation
Rethinking IRSTD: Single-Point Supervision Guided Encoder-only Framework is Enough for Infrared Small Target Detection
Towards Unsupervised Multi-modal Semantic Segmentation
XSemanticFlow: Cross Object Semantic Alignment for Zero-shot Manipulation
FD²: A Dedicated Framework for Fine-Grained Dataset Distillation
DINOde: Continuous Vision-Text Alignment for Open-Vocabulary Semantic Segmentation
Safe Generalization: Mitigating Catastrophic Forgetting in Single-Source Multi-Organ Segmentation via Collaborative Causal Learning
DecepGPT: Schema-Driven Deception Detection with Multicultural Datasets and Robust Multimodal Learning
Multi-scale Object-Aware Gaze Estimation via Geometric Reasoning
Learning Accurate Segmentation Purely from Self-Supervision
DetPO: In-Context Learning with Multi-Modal LLMs for Few-Shot Object Detection
Hierarchical Style Aggregation for Versatile Chinese Handwriting Generation
Fragmented Text Is Insufficient for Image Representation: Fine-Grained Correspondence in Multimodal Dataset Distillation
RA-SOD: Reliability-Aware RGB-T Salient Object Detection under Modality Degradation
NegAS: Negative Label Guided Attention and Scoring for Out-of-Distribution Object Detection with Vision-Language Models
Recurrent Cross-View Object Geo-Localization
MoMCE: Mixture of Modality and Cue Experts for Multimodal Deception Detection
Prototype-Conditioned Imagination for Compositional Zero-Shot Learning
Open-Vocabulary 3D Object Detection with Co-Distillation Discovery and Dual Guidance Robust Training
When W4A4 Breaks Camouflaged Object Detection: Token-Group Dual-Constraint Activation Quantization
VCP-DCN: Beyond Visual Concealed Property via Depth Collaborative Network for Camouflaged Object Detection
Maximum Spanning Tree Guided Confidence and Sparse Graph for Robust Noisy Label Learning
BIP: Bi-level Information Transfer and Completion Prompting for Visual Recognition with Missing Modalities
Geometric Regularization for Long-Tailed Semi-Supervised Learning via Gaussian Feature Bridges
Preserving Knowledge across Space and Time for Continual Video Deepfake Detection
ModuSeg: Decoupling Object Discovery and Semantic Retrieval for Training-Free Weakly Supervised Segmentation
Amplify, Aggregate, and Adjust: VideoMAE-based Holistic-Subtle Aggregation for Micro-Action Recognition
Local-to-global Cross-modal Coordination for Self-supervised RGB-T Tracking
Path-JEPA: Path Signature Based Predictive Learning for Skeleton Action Recognition
OP3DSG: Open-vocabulary Part-aware 3D Scene Graph Generation for Real-world Environments
Mitigate Modality-Asymmetric Forgetting via Stabilizing Visual Representations in CLIP-Based Class-Incremental Learning
Towards Sparsely Annotated Open World Object Detection
MAVFusion: Efficient Infrared and Visible Video Fusion via Motion-Aware Sparse Interaction
Towards Open-World Referring Expression Comprehension: A Benchmark with Training-free Multi-task Consistency Checker
Modality-Aware Out-of-Distribution Detection for Multi-Modal Action Recognition
Domain Arithmetic: One-Shot VLA Adaptation under Environmental Shifts
Mask-guided Semantic Alignment: Robust Learning with Noisy Labels via Temporal Attention Stability
Rethinking Cross-Spectral Image Generation via Shared-Specific Representation
FSDC-DETR: A Frequency-Spatial Domain Collaborative DETR for Small-Object Detection
Purify then Guide: Rethinking Domain Generalization for Multimodal Face Anti-Spoofing
MoAKE: Toward Unified All-in-One Action Quality Assessment via Mixture of Action Knowledge Experts
ASTAD: Asymmetric Style Transfer for Synthetic-to-Real Adaptation in Autonomous Driving
Warp-free Cross-view Geo-localization via Feature-space Consensus Mining
GenHOI: Generalized Hand-Object Pose Estimation with Occlusion Awareness
Rethinking Pseudo-Labels: Multi-Granularity Supervision for Domain Adaptive Object Detection
Group3D: MLLM-Guided Semantic Grouping for Open-Vocabulary 3D Object Detection
QST-SAM: Leveraging Cross-modal Instructions for Few-shot Referring Video Object Segmentation
ReliefSAM: A Geometry-Augmented Multi-Prior Adapter for Bas-Relief Segmentation
Beyond Isolated Objects: Relationship-aware Open Vocabulary Scene Understanding via 3D Scene Graph Analysis
GeoEdit: Geometry-Aware Object Editing via Dual-Branch Denoising
QASA: Quality-Guided K-Adaptive Slot Attention for Unsupervised Object-Centric Learning
RAG-3DSG: Enhancing 3D Scene Graphs with Re-Shot Guided Retrieval-Augmented Generation
Progressive Representation Learning for Multimodal Sentiment Analysis with Incomplete Modalities
AquaStereo: Enabling Underwater Stereo Matching via Depth-Conditioned Diffusion and Geometry Self-Distillation
Diffusion Model as a Generalized Segmentation Learner
SWSL: Semantic-aware Weakly Supervised Learning for 3D Motion Generation using 2D Motion Data
Liquid Fusion of Heterogeneous Representations Towards General Salient Object Detection
Beyond Categorical Matching: Intra-Class Graded Relevance Estimation for Cross-Modal 3D Retrieval
SOVTrack: Open-Vocabulary Multi-Object Tracking with Self-Supervised Pseudo Labeling and Feature Distillation
Iterative Refinement of Semantic and Spatial Representations for Open-Vocabulary Camouflaged Object Segmentation
HER-Count: Learning Hyper-Exemplar Representation for Generalized Zero-Shot Object Counting
FST-SAM3: Taming SAM~3 with Frequency-Spatio-Temporal Refinement for Video Polyp Segmentation
Revisiting Deepfake Detection: BCNet for Robust Generalization Beyond Semantic Dependence
REAL-OW: Rehearsal-free Open World Object Detection with Low-Rank Adaptation and Dual-Stage Objectness Modeling
Virtual Category-Guided Continual Generalized Category Discovery
CaPCL: Caption-Preserved Continual Learning for Text-to-Image Retrieval
Label-Free Text Prototype Adaptation for Open Vocabulary Segmentation
Toward Robust In-Context Segmentation via Concept Guidance
Aligning Anything: Hierarchical Motion Estimation for Video Frame Interpolation
Degradation-Robust and Temporally Consistent Infrared–Visible Video Fusion via One-step Diffusion Framework
CascadeProto: Cascaded Cross-Modal Prototype Purification via Entropy-Aware Learning for Few-Shot 3D Point Cloud Segmentation
Proposal Score Realignment Guided by Semantic Completeness for Weakly Supervised Temporal Action Localization
IP-SAM: Rethinking Prompt-Conditioned Segmentation for Prompt-Absent Deployment
DETR is Secretly a Multispectral Detector: Zero-Parameter Adaptation via Semantic Alignment
Debiased Textual Prompt Tuning for Enhancing Unknown Class Discovery
Context-Interactive Reasoning for Group Activity Detection
NegROI: Click-Centric Uncertainty-Guided Refinement with Scene-Conditioned Negative Prompts for Robust Interactive 3D Segmentation
DE2TR: Dual Evidence Detection Transformer for Video Temporal Grounding
CLIP-AUTT: Test-Time Personalization with Action Unit Prompting for Fine-Grained Video Emotion Recognition
Domain Adaptation with Adaptive Imagination for Visual Reinforcement Learning under Limited Target Data
P²Fusion: Prompt-based Progressive Infrared-Visible Image Fusion via Dual-Prior Distillation
DSeq-JEPA: Discriminative Sequential Joint-Embedding Predictive Architecture
Training-free Discriminative Patch Mining for Robust Few-Shot Recognition with CLIP
PASR: Pattern-Aware Scene-Conditioned Reasoning for Camouflaged Object Detection
μFlow: Leveraging Average Images for Improving Generalisation of Deepfake Faces Detectors
ORACLE-3D: Open-world Region-aligned Cross-modal Learning for Label-efficient 3D Scene Understanding
CURE: Contextual Debiasing and Unbiased Refinement for Training-Free Open-Vocabulary Semantic Segmentation
Unleashing Multimodal Large Language Models for Training-free HOI Detection in the Wild
Mitigating Pose–Scale Discrepancy Bias and Reforming Multi-Support Reasoning for Few-Shot Semantic Segmentation
CRISP: Calibration-Aware Visual State Space Duality for Remote Sensing Image Segmentation
Free-Lunch Augmentation by Revisiting Diffusion-Based Data Generation for Cross-Domain Few-Shot Object Detection
Breaking the Model Forgetting Cycle in Long-Incremental 3D Object Detection
Fourier Self-Supervision for Fine-Grained Generalized Category Discovery
GhostPoint: Self-Supervised Representation Learning by Hallucinating Occluded LiDAR Structure
Calibrate Before Adapt: Training-Free Pseudo-Label Calibration for Semi-Supervised Cross-Domain Few-Shot Detection
Learning Structurally Consistent Representations for Multi-View Radar Semantic Segmentation
Capturing Spectral and Spatial Patterns for Federated Remote Sensing Segmentation
Noise-Robust Facial Expression Recognition via Mamba-driven Neighbor Weight Refinement
Learning to Attract and Repel: Dual Quality Margin Learning for Face Recognition (DQM-Face)
General Incomplete Multimodal Learning via Dynamic Quality Perception
From Local Geometry to Global Pseudo-Labeling for Robust Positive–Unlabeled Learning under Covariate Shift
SCDL: Synergistic Confidence-Dispersion Learning for Semi-Supervised Video Polyp Segmentation
Defect-aware Hybrid Prompt Optimization for Zero-Shot Multi-type Anomaly Detection and Segmentation
Bayesian Uncertainty Attribution-Guided Fine-Tuning for Open-Set Action Recognition
MoE-KD: Your Teacher Model is Worth Mixture-of-Experts for Knowledge Distillation
Dual Distribution Estimation for Zero-shot Noisy Test-Time Adaptation with VLMs
HEM: a margin-based loss for visual categorisation tasks
The Label Imitation Game: Turing Test Network for Zero-Shot Pseudo-Label Pruning
FaceMoE: Mixture of Experts for Low-Resolution Face Recognition
An Inverse-Adversarial and Difficulty-Adaptive Robust Vision-Language Model
Orthogonal Knowledge Refreshing for Domain-Incremental Object Detection
Graph Coloring for Multi-Task Learning
S2-FracMix: Self-Saliency Fractal Mixup
FedNASP: Federated Vision-Language Navigation with Adaptive Step-wise Personalization
OrthoTailor: Geometric Orthogonalization for Conflict-Free Unified Fashion Generation
MixCompress: Mixture of Experts for Variable Rate Learned Image Compression
Path-level Hindsight Instructions for Semantic Exploration in Vision-Language Navigation
SVG-EAR: Parameter-Free Linear Compensation for Sparse Video Generation via Error-aware Routing
Learning to Recover Task Experts from a Multi-Task Merged Model
Tri-Efficient Transfer Learning for Point Cloud Videos
Focusing on What Matters: Saliency-Harnessing Accurate Routing for Diffusion MoE
Residual-Guided Expert Specialization for Incomplete Multimodal Learning
Unified Multi-Layer Subspace Modeling for Cross-Domain OOD Detection
GCMRD: Global Consistency Multi-teacher Robustness Distillation
Distill Once, Adapt Life-Long: Exploring Dataset Distillation for Continual Test-Time Adaptation
DriveFine: Refining-Augmented Masked Diffusion VLA for Accurate and Robust Driving
Prefill-Time Interventions against Adversarial Attacks on Large Vision-Language Models
Benchmarking Federated Learning & Knowledge Distillation for Point Cloud Classification
CL-Anomaly: Layer-Adaptive Mixture-of-Experts with Multimodal Large Language Model for Continual Learning in Anomaly Detection
Towards Purified Multi-Label Test-Time Adaptation of Vision-Language Models
Holistic Optimal Label Selection for Robust Prompt Learning under Partial Labels
Rethinking Adversary in Semantic Segmentation: An Out-of-Distribution Perspective
RT-DETRv4: Painlessly Furthering Real-Time Object Detection with Vision Foundation Models
Dive into the implicit biases of low-rank vision-language alignment
CollectionLoRA: Collecting 50 Effects in 1 LoRA for Deployment
VisNec: Measuring and Leveraging Visual Necessity for Multimodal Instruction Tuning
Learning 1-Bit LiDAR-based Localization with Auxiliary Objective
Task-Agnostic Incremental Vision-Language Object Detection via Prompt Augmentation and Distribution-Aware Fusion
Point Ladder Tuning: Parameter-Efficient Hierarchical Adaptation for 3D Point Cloud Understanding
SLER-IR: Spherical Layer-wise Expert Routing for All-in-One Image Restoration
AnchorGUI: Asymmetric Memory for Dual-Scale Learning in GUI Navigation
Prevention over Correction: Learning Aligned Representations in One-shot Federated Learning
TrustCLIP: Learning Private Visual Features via Adversarial Reconstruction
MED-LCDS: Multi-Expert-Domain CLIP Classification via Logit Calibration
Training-Free Task Classification for Multi-Task Model Merging
Rank-Aware Hyperbolic Alignment for Vision–Language Dataset Distillation
Attention Misses Visual Risk: Risk-Adaptive Steering for Multimodal Safety Alignment
SegPAR: Class-Centric Decision-Based Sparse Attack for Semantic Segmentation
SUM: Unified Geometric Surgery on Spatio-Temporal Adaptation Vectors for Federated Class Incremental Learning
FeDepth: Federated Learning for Depth Estimation under Robot Heterogeneity
DualTAP: A Dual-Task Adversarial Protector for Mobile MLLM Agents
Distill on a Diet: Efficient Knowledge Distillation via Learnable Data Pruning
Why Feature Magnitude Deceives OOD Detectors: An Angular Separation Perspective
SIMPLER: Efficient Foundation Model Adaptation via Similarity-Guided Layer Pruning for Earth Observation
Fisher-Routed Mixture of Experts for Federated Class-Incremental Learning
Comprehensive Robustness Analysis of LiDAR-based 3D Object Detection in Autonomous Driving
SCoT: Similarity-guided Conflict-aware Task Consolidation for Continual VQA
Zero-Shot Quantization for Object Detectors using Off-the-Shelf Generative Models
DA-MergeLoRA: Hypernetwork-Based LoRA Merging for Few-Shot Test-Time Domain Adaptation
Importance-Aware Low-Rank Distillation of Diffusion Transformers
BATQuant: Outlier-Resilient MXFP4 Quantization via Learnable Block-wise Optimization
Molecular Identifier Visual Prompting and Verifiable Reinforcement Learning for Chemical Reaction Diagram Parsing
Prototype Normalization: Optimizing Prototype Separation for Heterogeneous Federated Learning
ReTarget: Representation Transformation via Adversarial Regularization for Geometric Misalignment
Flash-DD: An Ultra Parameter-Efficient Approach to Dataset Distillation
DroneFINE: Domain-Aware Parameter-Efficient Fine-Tuning of Vision-Language Detectors for Drone Images
MUSE: Unlocking Timestep as Native Task Steering for One-Step Dense Prediction
RUTaL: Residual Upcycling with Task Ladder for Efficient Multi-Task Learning
Learning from Reliable Negatives: Confidence-Anchored Test-Time Adaptation for GUI Grounding
Neutralizing Token Aggregation via Information Augmentation for Efficient Test-Time Adaptation
SPDA: Efficient Online Test-Time Adaptation for Promptable Medical Segmentation
Unified Prediction and Planning via Conflict-Aware Disjoint Parameter Training
One Trap to Block Them All: Defending Encoder Stealing via Isotropic Uniformity
Revisiting Parameter Redundancy in Vision-Language-Action Models: Insights from VLM-to-VLA Adaptation
TextDS: Parameter-Efficient Representation Alignment for Scene Text Detection under Distribution Shifts
Interference-Aware Continual Vision–Language Learning via Instance-Level Expert Routing
Diffusion to Obfuscation: Time-Adaptive Synthesized Generation Against Gradient Leakage Attacks in Federated Learning
Identifiable Gated Residual Personalization for Federated Parameter-Efficient Fine-Tuning
TSEmbed: Unlocking Task Scaling in Universal Multimodal Embeddings
DiscoVL: Unveiling Disentangled Cross-Modal Representation Learning via Orthogonal Adversarial Regularization for Vision-Language Models
Curvature-Guided Mixing for MLLM Adaptation
PACO: Stabilizing Vision Embeddings along Local Paths for Robust Vision-Language Models
Aggregating Cross-Domain Knowledge via Learnable Tokens for Multi-Teacher Distillation
TIR-Agent: Training an Explorative and Efficient Agent for Image Restoration
Multi-Scale Representation Alignment for Visual Autoregressive Modeling with Mixture of Experts
R2M: Real-Aware Residual Model Merging for Robust and Generalizable Deepfake Detection
BioMedVR: Confusion-Aware Mixture-of-Prompt Experts for Biomedical Visual Reprogramming
Robust Trajectory Distillation: Hybrid Reweighting Meets Teacher-Inspired Targets
ViQ: Text-Aligned Visual Quantized Representations at Any Resolution
SiGMA: Sign-Guided Merging and Adaptation framework for Multimodal Continual Instruction Tuning
COVERT: Privacy-Preserving Covariant Obfuscation for VLMaaS via Exact Reparameterization and Tailored Tuning
Symbiotic-MoE: Unlocking the Synergy between Generation and Understanding
VD-LoRA: Adaptive Reuse of Low-Rank Directions for Continual Learning
Going Deep: Deep Visual Prompting with LoTeP
Probe, Anchor, and Amend: Active Test-Time Adaptation of Vision-Language Models
Robust onion: Peeling Open Vocab Object Detectors Under Noise
Foundation Model Selection for Remote Sensing via a Constraint-Aware Agent
Multi-Hypothesis Test-Time Adaptation to Mitigate Underspecification
GeMoE: Gating Entropy is All You Need for Uncertainty-aware Adaptive Routing in MoE-based Large Vision-Language Models
LDC-MTL: Balancing Multi-Task Learning through Scalable Loss Discrepancy Control
Locality-Aware Continual Unlearning for Diffusion Models
Exploiting Local Flatness for Efficient Out-of-Distribution Detection
FedDO: Dynamic Client Optimization for Adaptive Federated Learning
Indelible Backdoors: On the Limits of Post-Training Defenses
Attention-Logit Steering to Compositional Generalization for Continual VQA
Condensing Large-Scale Datasets Directly with Minimal Information Loss
Isotropic Embedding Perturbations for Robust Vision Language Encoders
MoE-ViE: Mixture of Experts Vision Encoder for Efficient Image and Video Understanding
Stabilizing Ultra-Low-Bit Quantization of Multimodal LLMs via Global Bit Allocation
Token-level Response-visual Attention Guidance for Multimodal LLMs Knowledge Distillation
Prune Once: Retraining-Free Task-Agnostic Pruning for Vision-Language Models
KATANA: Knowledge-Aligned Topology-Aware Neural Agents for RL-Driven Vision-Language Model Compression
Breaking Rigidity in Adversarial Patch Attacks
LANCE: Low Rank Activation Compression for Efficient On-Device Continual Learning
Low-Rank Ternary Adaptation for Fine-Tuning Transformers
GeoWorld: Providing Full-frame Geometry Features to Facilitate 3D Scene Generation
Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding
Argus: Metric Panoramic 3D Reconstruction for Indoor Scenes
LiFlow: Flow Matching for 3D LiDAR Scene Completion
Event-LiDAR: 3D Eventification for Efficient Point Cloud Processing
CanoVerse: 3D Object Scalable Canonicalization and Dataset for Generation and Pose
OCA: ODE-Driven Cross-Attention for Image-to-Point-Cloud Registration
SuperFlex: Deformable Superquadrics for Point Cloud Decomposition
FDR-Occ: Factorized Dense Routing for Full-Spectrum 3D Occupancy Prediction
MAGiSt3R: Multi-Agent Feed-forward 3D Reconstruction from Monocular RGB Videos
Map-Det3D: Metric Feed-Forward 3D Reconstruction Prior for Multi-view 3D Object Detection from Streaming Inputs
GraphCPD: Coherent Point Drift for Point Cloud Registration via Graph Signal Processing
LangLoc: “Tell Me What You See”
Unsupervised Point Cloud Registration via Training-Time Semantic Guidance
MSVS-VAE: Multi-Scale Anchored VecSet for High-Fidelity 3D Reconstruction
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding
DepWorldSG: Depth-Aware 3D Semantic Scene Graph Generation via World-Model Priors
PGCR: Pose–Geometry Coupled Reasoning for Image-to-Point Cloud Registration
CasaMaestro: Multi-View Panoramas for House-Scale 3D Reconstruction
SynCity 3000: Bootstrapping Scene-Scale 3D Diffusion
RGB-Pointmap Pretraining for Unified 3D Scene Understanding
Phase-Aligned RoPE for Mixed-Resolution Diffusion Transformer
YouTube-Occ: Learning Indoor 3D Semantic Occupancy Prediction from YouTube Videos
Ex-Sim(3)-Reg: 2D-3D Correspondence Pruning via Extended Sim(3) Registration
Attention-DP3: Spatially Object-aware 3D Diffusion Policy via Geometry-aligned Attentional Conditioning
Hierarchical and Holistic Open-Vocabulary Functional 3D Scene Graphs for Indoor Spaces
Pose Anything Anywhere: Model-free Object Poses from Arbitrary References
HHA: Hierarchical Hyperbolic Constraints for Imperceptible Point Cloud Attacks
SuperVoxelGPT: Adaptive and Ordered 3D Tokenization for Autoregressive Shape Generation
Fixed Reality, Diffused Possibility: Disentangling Stochastic and Deterministic Latent for Cluttered Grasping
TabletopGen: Tabletop Scene Generation and Interactive Simulation for Robotic Manipulation
GridFlow: Structured Latent Flow for Seamless City-Scale 3D Point Cloud Generation
SNOW: Spatio-Temporal Scene Understanding with World Knowledge for Open-World Embodied Reasoning
E-M3RF: An Equivariant Multimodal 3D Re-assembly Framework
UniPart: Towards Zero-shot Language-Grounded 3D Part Segmentation for Embodied Interaction
Seen2Scene: Completing Realistic 3D Scenes with Visibility-Guided Flow
WorldFlow3D: Flowing Through 3D Distributions for Unbounded World Generation
SceneOrchestra: Efficient Agentic 3D Scene Synthesis via Full Tool-Call Trajectory Generation
RayMap3R: Inference-Time RayMap for Dynamic 3D Reconstruction
RegVGGT: Sustainable Visual Geometry Grounding for Streaming via Regulated Memory
Scal3R: Learning Efficient Multi-Relative Pose Query for Scalable Online 3D Reconstruction
GaussianGPT: Towards Autoregressive 3D Gaussian Scene Generation
On Geometric Understanding and Learned Priors in Feed-forward 3D Reconstruction Models
TASE: Truncation-Aware Semantic Embeddings for 3D Scene Understanding and Editing
Revisiting Scene Graph Generation from the Perspective of Detector-Conditioned Reachability
UniQueR: Unified Query-based Feedforward 3D Reconstruction
SEM-ROVER: Semantic Voxel-Guided Diffusion for Large-Scale Driving Scene Generation
LESV:Language Embedded Sparse Voxel Fusion for Open-Vocabulary 3D Scene Understanding
PUF: Plug-and-Play Uncertainty-Aware Fusion for Online 3D Scene Graph Generation
Rethinking Training and Inference for Trajectory Forecasting: Linking Winner-Take-All back to GMMs
Spatial-TTT: Streaming Visual-based Spatial Intelligence with Test-Time Training
Zero-Shot Novel Depth Synthesis Using Foundation Models Scene Representations
SLAM-Former: Putting SLAM into One Transformer
Pixel-wise Geo-registration of Drone and Satellite Images
MessyKitchens: Contact-rich object-level 3D scene reconstruction
Occlusion-Resilient Category-Agnostic Pose Estimation with Conditional Flow Matching
Abstract the Layout, Focus the Detail: A Dual-Granularity Representation Framework for Zero-Shot 3D Visual Grounding
SEAR: Simple and Efficient Adaptation of Visual Geometric Transformers for RGB+Thermal 3D Reconstruction
Sparse-Aware Vector Quantization for Bandwidth-Efficient Collaborative 3D Semantic Occupancy Prediction
Point Diffusion Mamba: Unified Diffusion-State-Space Modeling for Single-View 3D Reconstruction under Data Scarcity
SegFly: A 2D-3D-2D Paradigm for Aerial RGB-Thermal Semantic Segmentation at Scale
Unfold The World: Factorize 4D Properties in Reinforcing Spatial Understanding
Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding
A First Exploration of Neuromorphic OT-CFM for Multi-Speaker VSR
ReGRPO: Reflection-Augmented Policy Optimization for Tool-Using Agents
Towards Interactive Global Geolocation Assistant
Vero: Open Reinforcement Learning Recipes for Visual Reasoning
SWIFT: Spatial-Window Integrated Frequency-aware Token Pruning for Efficient MLLMs on Edge Devices
PhysAlign: Learning Physical Priors for Dynamical Event-Driven Video Generation via Representation Alignment
MarineEVT: Advancing Event-Centric Marine Video Understanding via Visual Tool Reasoning
SpaMEM: Benchmarking Dynamic Spatial Reasoning via Perception–Memory Integration in Embodied Environments
LASER: A Corrective Lens for LVLMs via Visual Attention Preservation and Sink Suppression
FlowCIR: Semantic Transport via Flow Matching for Zero-Shot Composed Image Retrieval
OmniCoT: A Benchmark for Global and Multi-Step Panoramic Reasoning
Diagram-MMU: A Multi-Modal Benchmark for Scientific Diagrams
Towards Spatial Trace with Reasoning in Vision-Language Models for Robotics
MG2-RAG: Multi-Granularity Graph for Multimodal Retrieval-Augmented Generation
AMCI: Unlock the Potential of Large Multimodal Models for Fine-grained Open-world Classification via Adaptive Memory Context Injection
OmniSch: A Multimodal PCB Schematic Benchmark For Structured Diagram Visual Reasoning
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models
Embed-RL: Reinforcement Learning for Reasoning-Driven Multimodal Embeddings
ATP-Bench: Towards the Agentic Tool Planning for MLLM Interleaved Generation
Combating Textual Noise and Redundancy: Entropy-Aware Dense Visual Token Pruning
Data-Free Client Contribution Estimation via Logit Maximization for Federated Learning
SyntheticDoc: A Large Synthetic Dataset for Document Unwarping and Illumination Correction
Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning
Spectral Evolution-Guided Token Pruning in Large Multimodal Models
AnyGround3D: Towards Grounding Any 3D Object in the Wild via 2D-to-3D Lifting
Dense Video Understanding with Inter-tokenization Acceleration
HERO: Enhancing Multimodal Faithfulness via Dynamic Entropy-Aware Reinforcement Learning
Video-Text Alignment Model for Sign Language Translation
From Hallucination to Grounding: Diagnosing Visual Spatial Intelligence via CRISP
OneWorld: Taming Scene Generation with 3D Unified Representation Autoencoder
Enhancing Interpretability in CLIP with Optimal Transport-based Submodular Optimization for Ophthalmic Imaging
Semantic Generative Tuning for Unified Multimodal Models
MultiHaystack: Benchmarking Multimodal Retrieval and Reasoning over 40K Images, Videos, and Documents
Counting Trees from Satellite Imagery with Noisy Supervision
Towards Spatial Supersensing in the Wild
PanoRec: Spatially-Structured Sequence Modeling for Multi-Granularity Panoramic Retrieval
LoGeR: Long-Context Geometric Reconstruction with Hybrid Memory
ForeSea: AI Forensic Search with Multi-modal Queries for Video Surveillance
Toward Interpretable Analysis of Whole-slide Pathology Images via Large Language Model-based Agentic Reasoning
GeoSolver: Scaling Test-Time Reasoning in Remote Sensing with Fine-Grained Process Supervision
UBone3D: Physics-Rectified Conditional Flow Matching for Anatomical 3D Shape Completion from Ultrasound
Lumina-OmniLV: A Unified Multimodal Framework for General Low-Level Vision
HippoCamp: Benchmarking Contextual Agents on Personal Computers
CoCo-IR: Conversational Composed Image Retrieval
LiveK12Bench: Have Large Multimodal Models Truly Conquered High School-level Examinations?
Staying VIGILant: Mitigating Visual Laziness in MLLMs via Information-Theoretic Alignment
Co-Steer: Cross-Modal Collaborative Steering for Jailbreaking MLLMs
NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding
GradingBench: Evaluating End-to-End Compositional Reasoning of MLLMs for Automated Exam Grading
AnchorPrune: Relevance-Anchored Contextual Expansion for Visual Token Pruning
AnE: Pushing the Reasoning Frontier of Multimodal LLMs via Anchor Evolution
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs
EndoCoT: Scaling Endogenous Chain-of-Thought Reasoning in Diffusion Models
When Thinking Hurts: Mitigating Visual Forgetting in Video Reasoning via Frame Repetition
Parallel Vision Token Scheduling for Fast and Accurate Multimodal LMMs Inference
Knowledge-Centric Agents for Workflow Generation in ComfyUI
TIR-Bench: A Comprehensive Benchmark for Agentic Thinking-with-Images Reasoning
Learning Sample-wise Rank-Aware Interpolation Weights for Composed Visual Data Retrieval
Beyond Description: Cognitively Benchmarking Fine-Grained Action for Embodied Agents
IQA-T1: Tool-based Visual Evidence Reasoning for Image Quality Assessment
OmniMapBench: Benchmarking Visual-Centric Reasoning on Diverse Map Documents
Keeping the Evidence Chain: Semantic Evidence Allocation for Training-Free Token Pruning in Video Temporal Grounding
ConsiSpace: Learning Geometric Consistency Matters for Video Spatial Reasoning
Goku: A Million-Scale Universal Dataset and Benchmark for Instruction-Based Video Editing
Mixture of Specialized Vision Experts: Unlocking Complementary Visual Insights for Faithful MLLM Reasoning
AdaThinking-E: One-Token Entropy Regulation for Adaptive Thinking
Constructing and Interpreting Digital Twin Representations for Visual Reasoning via Large Language Models and Reinforcement Learning
MindBlock: Probing Spatial Assembly and Structure in Unified Multimodal Models
VERDICT: Training-Free Step-Wise Verification of Multimodal Reasoning via Disagreement-Aware Consensus
EGVLR: Evidence-Grounded Vision–Language Reinforcement for Anomaly Reasoning
BrepCoder: A Unified Multimodal Large Language Model for Multi-task B-rep Reasoning
Unbalanced Optimal Transport for Efficient Visual Document Retrieval
Together, Then Apart: Balancing Alignment and Distinctiveness for Multimodal Survival Analysis
DocLayout-VL: A Foundational Model for Hierarchical, Open-set, and Promptable Document Layout Segmentation
FlashVLM: Text-Guided Visual Token Selection for Large Multimodal Models
MotionChain: Fine-Grained Video Motion Understanding via Structured Decomposition
ToDRE: Effective Visual Token Pruning via Token Diversity and Task Relevance
LightSTAR: Efficient Visual Document Retrieval via Lightweight Selection with Vision-Adaptive Refinement
Generalize LMMs to Versatile Visual Modalities via Fabricated Modality Synthesis
ViewFusion: Structured Spatial Thinking Chains for Multi-View Reasoning
MotionAtlas: A High-Quality Dataset and Benchmark for Dense Motion Captioning
Improving Reasoning in Vision-Language Models via Perception Verified Self-Training
VisCritic: Visual State Comparison as Process Reward for GUI Agents
Spotlight: Identifying and Localizing Video Generation Errors Using VLMs
SVGEval: A Vision-Grounded Framework for Perceptual-Quality Benchmarking and Evaluation in Text-to-SVG Generation
Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe Extraction
ADAPT: Attention Dynamics Alignment with Preference Tuning for Faithful MLLMs
See & Sniff: Learning Visuo-Olfactory Representations
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs
Multi-label Instance-level Generalised Visual Grounding in Agriculture
CRAG-MM-Diagnostics: Enabling Stage-Wise Analysis of Knowledge-Intensive VQA
Memory Tree Guided Key Frame Querying for Efficient 3D Question Answering
Personalizing MLLMs via Reinforced Multimodal Reference Game
Global Logic and Local Search: Dual-Stream Multimodal In-Context Learning for Verifiable Industrial Anomaly Detection
Global-Local Dual Perception for MLLMs in High-Resolution Text-Rich Image Translation
P-MTP: Efficient Document Parsing via Multi-Token Prediction with Progressive Depth Scaling
HSD: Training-Free Acceleration for Document Parsing Vision-Language Models with Hierarchical Speculative Decoding
SAFE-EQA: Semantic-Aware Efficient Exploration for Embodied Question Answering
CMDR: Contextual Multimodal Document Retrieval
RoadBench: Benchmarking MLLMs on Fine-Grained Spatial Understanding and Reasoning under Urban Road Scenarios
RegRet: Enhancing Region-Level Retrieval in Large Multimodal Models
The Illusion of High Utility in Safety Alignment of Text-to-Image Diffusion Models
COMPASS: Grounding Composition-Intent Guidance in Unified Multimodal Models
From Illusion to Intention: Visual Rationale Learning for Reliable Evidence Acquisition
Falcon: Functional Assembly and Language for Compositional Reasoning in X-ray
On the Faithfulness of Post-Hoc Concept Bottleneck Models
NAPA: Natively Multimodal Autoregressive Perception Architecture
Information-Regularized Attention for Visual-Centric Reasoning
Eliciting Self-Verification in Multimodal Reasoning Agents with Reinforcement Learning
Why and When Visual Token Pruning Fails? A Study on Relevant Visual Information Shift in MLLMs Decoding
Mitigating Sycophancy in Multimodal Chart Understanding via Vision-Grounded Verification
Beyond Where to Look: Trajectory-Guided Reinforcement Learning for Multimodal RLVR
Transport Discrepancy as a Reliability Signal for Vision-Language-Action Models
EpiBench: Benchmarking Multi-turn Research Workflows for Multimodal Agents
Wake up for Touch! Mask-isolated Tactile Alignment Learning in MLLMs
ProMSA:Progressive Multimodal Search Agents for Knowledge-Based Visual Question Answering
Wavelet-based Intra-video Counterfactual Reasoning for Video Question Grounding
How Far Are Video Models from True Multimodal Reasoning?
MV-STRIDE: Enabling MLLMs to Master Multi-View Spatial Reasoning via Hierarchical Capability Modeling
From Reasoning to Pixels: Grounded Medical Multimodal LLMs for VQA and Segmentation
Real5-OmniDocBench: A Full-Scale Physical Reconstruction Benchmark for Robust Document Parsing in the Wild
GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding
BaFCo: A Document Understanding Benchmark for Complex Bangla Form Comprehension
Iterative Perceptual Alignment for VLMs via Deterministic Reconstruction Feedback
HART: High-Resolution Annotation-Free Reasoning Technique through a Closed-loop Framework
Caption Bottleneck Models
CogniCred: A Dataset and Benchmark for Cognitive Credential Forgery Detection
Error-Driven Scene Editing for 3D Grounding in Large Language Models
GridVQA-X: A Diagnostic Framework for Evaluating Multimodal Explainability Methods
Geo-DPO: Aligning Semantic Intent with Geometry for 3D Affordance Segmentation
CiQi-Agent: Aligning Vision, Tools and Aesthetics in Multimodal Agent for Cultural Reasoning on Chinese Porcelains
ViTAL‑X: Video-Text Alignment with Cross‑Modal Temporal Edits
Learning Spectral and Polarimetric Clues for One-to-Multimodal Novel View Synthesis
BBQ-V: Benchmarking Visual Stereotype Bias in Large Multimodal Models
LatentPilot: Scene-Aware Vision-and-Language Navigation by Dreaming Ahead with Latent Visual Reasoning
OmniNWM: Unifying the State-Action-Reward Triad for Closed-Loop Panoramic Driving Navigation World Models
ChronoFlow Policy: Unifying Past-Future Interaction Flow in Visuomotor Policy Learning
DriveVA: Video Action Models are Zero-Shot Drivers
RoMan-4D: Learning Robot Arm Manipulation from 4D World Models
OmniFit: Multi-modal 3D Body Fitting via Scale-agnostic Dense Landmark Prediction
Panoramic Affordance Prediction
HSImul3R: Physics-in-the-Loop Reconstruction of Simulation-Ready Human–Scene Interactions
CabinSI: Omni-Cabin Spatial Reasoning through Explicit Visual Cognitive Maps
Towards Generalizable Robotic Manipulation in Dynamic Environments
DriveWeaver: Point-Conditioned Video Inpainting for Controllable Vehicle Insertion in Autonomous Driving Simulation
OmniStream: Mastering Perception, Reconstruction and Action in Continuous Streams
ASSCG: Just-Right Gating over Chattering for Fast–Slow LLM Planning in Autonomous Driving
S-VAM: Shortcut Video-Action Model by Self-Distilling Geometric and Semantic Foresight
XYZ-IBD: Benchmarking Robust 6D Object Pose Estimation under Real-World Industrial Complexity
A scalar per patch from pre-trained ViTs enables fast moving navigation in the real world
ORION: Ordinal Neural Collapse as a Representation Prior for Visual Navigation
Ultra3D: Efficient and High-Fidelity 3D Generation with Part Attention
Grasp-Oriented Non-Prehensile Manipulation via Learning a Graspability Field
VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Model
CoMind: Understanding Collaborative Human Activity from Multiple Minds and Views
Learning Transferable Dynamics Priors from Action to World Modeling
CoMaTrack: Competitive Multi-Agent Game-Theoretic Tracking with Vision-Language-Action Models
Walk through Paintings : Ego-centric World models from Internet Priors
TEX-Drive: Temporal Perception Meets Experience-Guided Mixture-of-Experts for End-to-End Autonomous Driving
CausalDrive: Real-time Causal World Models for Autonomous Driving
Driver-WM: A Driver-Centric Traffic-Conditioned Latent World Model for In-Cabin Dynamics Rollout
VGGT-World: Transforming VGGT into an Autoregressive Geometry World Model
RoboStream: Weaving Spatio-Temporal Reasoning with Memory in Vision-Language Models for Robotics
QuantV2X: A Fully Quantized Multi-Agent System for Cooperative Perception
PixelPilot: Scalable Vision-Language-Action Models for End-to-End Autonomous Driving
PersonaDrive: Controllable Trajectory Prediction with Multi-Dimensional Driving Personas
CoGoal3D: Collaborative 3D Object Detection with 3D-Aware Fusion and Refinement
XDen-1K: A Density Field Dataset of Real-World Objects
NavWM: A Unified Navigation World Model for Foresight-Driven Planning
FlowDec: Temporal Conditional Flow Decorruptor for Robust Continuous Vision-Language Navigation
Vulnerability of Privacy-Preserving Visual Localization against Diffusion-based Attacks
Stabilizing Real-World Visual Active Tracking with Action-Smooth Test-Time Adaptation
RelAfford6D: Relational 6D Affordance Graphs for Constraint-Driven Robotic Manipulation
MobileManiBench: Simplifying Model Verification for Mobile Manipulation
Agentic Collaborative Cognition for Zero-Shot 3D Understanding
Deconfounded Lifelong Learning for Autonomous Driving via Dynamic Knowledge Spaces
DreamerAD: Efficient Reinforcement Learning via Latent World Model for Autonomous Driving
When the City Teaches the Car: Label-Free 3D Perception from Infrastructure
UniDrive-WM: Unified Understanding, Planning and Generation World Model For Autonomous Driving
BEVLM: Distilling Semantic Knowledge from LLMs into Bird's-Eye View Representations
StreamSpatial: A Benchmark and Framework for Streaming 3D Visual-Spatial Reasoning
Driving like yourself: A Benchmark for Closed-Loop Personalized End-to-End Autonomous Driving
RAF: Reliability-Aware Fusion of Camera, LiDAR, and 4D RADAR for Robust 3D Object Detection in Adverse Weather
PriorMaskMap: Robust Online Vectorized Map Construction with Biased Priors
UniTeD: Unified Temporal Diffusion for Joint Perception and Planning in Autonomous Driving
StructPolicy: Structure-Guided Imitation Learning Robust to Visual Domain Shifts
RadarGen: Automotive Radar Point Cloud Generation from Cameras
MA-VLA: Multi-Arm Vision-Language-Action Model for Collaboration and Compositional Generalization
ESCAPE: Episodic Spatial Memory and Adaptive Execution Policy for Long-Horizon Mobile Manipulation
World-in-Loop: Online Correction via Event-Triggered World Models for Robust VLA Policies
LSRM: High-Fidelity Object-Centric Reconstruction via Scaled Context Windows
SGC-Lane: Monocular 3D Lane Detection with Standard-Definition Map Guidance and Lane Completion
WALL-EVE: World Alignment with Rule Learning in Visual Environments
DiNBV-Grasp: Real-Time Distance-Aware Two-Stage Next-Best-View for Robotic Grasping
C2E: Boosting Ego-Only 3D Object Detection via Multi-Teacher Contrastive Knowledge Distillation
SP-TransientBench: A Real-Captured Single Photon Perception Benchmark
Guide, Think, Act: Interactive Embodied Reasoning for Vision-Language-Action Model
Noise is a Good Teacher: A Noise-Driven Framework for Robust Collaborative Perception
PhysMani: Physics-principled 3D World Model for Dynamic Object Manipulation
RESOLVE: A Multi-Resolution and Multi-Modal Dataset for Roadside Cooperative Perception
CoReLIN: Constraint-based Reasoning for Zero-shot Lifelong Interactive Navigation
E3VS-Bench: A Benchmark for Viewpoint-Dependent Active Perception in 3D Gaussian Splatting Scenes
DiverseAD: A Large-Scale Driving Dataset with Diverse Atmospheric Conditions
Sen-Cap: Sensor-Flexible and Noise-Resilient Human Motion Capture via LiDAR-Camera Integration
Markov-Renewal Single-Photon LiDAR Simulator
OpenSpatial: A Principled Data Engine for Empowering Spatial Intelligence
MolmoWeb: Open Visual Web Agent and Open Data for the Open Web
MM-Nav: Multi-View VLA Model for Robust Visual Navigation via Multi-Expert Learning
What if? Emulative Simulation with World Models for Situated Reasoning
TriO: Tri-Modal Unsupervised Occupancy World Model for Anything Perception
TAIHRI: Task-Aware 3D Human Keypoints Localization for Close-Range Human-Robot Interaction
SIMON: SImultaneous Multi-Object Navigation
3DWay: Generalizing Robot Manipulation via 3D Consistent Waypoints
Towards Unified World Models for Visual Navigation via Memory-Augmented Planning and Foresight
VVSim: A Large-Scale Aerial-Ground Dataset and Benchmark for Cooperative Perception
MindDrive: A Vision-Language-Action Model for Autonomous Driving via Online Reinforcement Learning
LEO-Fuse: A Modality- and Task-Agnostic Universal Framework for Multimodal Human Sensing
UniBYD: A Unified Framework for Learning Robotic Manipulation Across Embodiments Beyond Imitation of Human Demonstrations
OpenEarthAgent: A Unified Framework for Tool-Augmented Geospatial Agents
E-VLA: Event-Augmented Vision-Language-Action Model for Dark and Blurred Scenes
VOCA: Visual Odometry with Codec Awareness
LaGen: Towards Autoregressive LiDAR Scene Generation
Deform360: A Massive Multi-view Visuotactile Dataset for Deformable World Models
Two-Way Street: Efficient VSLAM using Collaborative In-Sensor and Off-Sensor processing
Boba: Batched Simulation for Physics-Based Gaussian Digital Twins
Lifting Ego World Models for Planning and Control
Sentinel: Embodied Cooperative Spatial Reasoning and Planning
FindingDory: A Benchmark to Evaluate Memory in Embodied Agents
RAE-NWM: Navigation World Model in Dense Visual Representation Space
WildCity: A Real-World Dataset for City-Scale Rendering and Beyond
VPA-WM: Vision-Priors-Aligned World Models for Robust Visual Reinforcement Learning
Self-Evolving Just-In-Time Memory for Proactive Embodied Safety
Drive2Danger: Deceive End-to-End Autonomous Driving with Risky Instance Recognition
STEP: Spatial Thinking and Egocentric Pointing for Embodied Instruction Following
Tac2Real: Reliable and GPU Visuotactile Simulation for Online Reinforcement Learning and Zero-shot Real-World Deployment
Hypothesis Graph Refinement: Hypothesis-Driven Exploration with Cascade Error Correction for Embodied Navigation
Pondering the Way: Spatial-perceiving World Action Model for Embodied Navigation
AMCoNav: Asynchronous Multi-module Collaborative Framework for Embodied Visual Navigation
LMGenDrive: Bridging Multimodal Understanding and Generative World Modeling for End-to-End Driving
NutriBench-Kitchen: Benchmarking Embodied AI for Nutrition Management
Generative Lane Topology Reasoning via Autoregressive Model with Geometry Prior
Grounding Sim-to-Real Generalization in Dexterous Manipulation: An Empirical Study with Vision-Language-Action Models
The Language of Visual Attention: Modeling Scanpaths via Autoregressive Token Prediction
WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation
EvoWorld: A World-Model-Centric Framework for Continuous Self-Evolution of Modular Embodied Skills
WildWorld: A Large-Scale Dataset for Action-Conditioned World Modeling with Explicit State Annotations
RGBT-GroundBench: Visual Grounding Beyond RGB in Complex Real-World Scenarios
MobileOcc: A Human-Aware Semantic Occupancy Dataset for Mobile Robots
SkillSpotter: Pose-Aware Multi-View Skilled Action Detection and Grading in Ego-Exo Videos
Stand Up and Move: Benchmarking Interactive Spatial Intelligence in WalkerBench
HERO: Heterogeneous Evidential Robust Object-Level Collaborative Perception
RoboTALES: Learning Reasoning-Guided Robot Policies via Task-Aligned Simulated Futures
MV2GF: Multi-view Pedestrian Detection with a Visual Geometric Foundation Model
Hi-Nav: Hierarchical Framework for Continuous Vision-Language Navigation via Map Guidance and Waypoint Reasoning
UECP: Uncertainty-Enhanced Collaborative Perception
World Models for Learning Dexterous Hand-Object Interactions from Human Videos
Hierarchical 3D Scene Graph Construction and Belief-based Planning for Semantic Navigation
Driving is a Game: Combining Planning and Prediction with Bayesian Iterative Best Response
Twin-DAgger: Synergizing Digital Twins and Human Corrections for Efficient Robot Manipulation
LiSTAR: Ray-Centric World Models for 4D LiDAR Sequences in Autonomous Driving
CausalVAE as a Plug-in for World Models: Towards Reliable Counterfactual Dynamics
Unpaired Geometry-Guided Sim2Real Translation for Autonomous Driving
SAFER-Activities: A Dataset for Smart Assessment of Fall Events and Routine Activities
Thinking from the Robot’s View: The CoT-HRC Benchmark for Human Intent Reasoning in Embodied Collaboration
Exploratory, Communicative, and Deployable: Vision-Driven Embodied Agents for Open-World Mobile Manipulation
CooperScene: Multi-Modal Cooperative Autonomy Benchmark with C-V2X Communication Characterization
Towards Metric-Agnostic Trajectory Forecasting
DH-VLM: Dual-Horizon Cooperative Latent Reasoning for Autonomous Driving
SynVAR: Synergizing Spatial and Semantic Alignment in Visual Autoregressive Model
Geometry-Aware Spatio-Temporal Context Modeling for 4D Occupancy Forecasting
One Demonstration Is Enough for Real-World Robotic Reinforcement Learning
RECO: Region-Aware Compensation for Extrinsic Perturbations in Roadside 3D Detection
AdaDexGrasp: Adaptive Dexterous Grasping via 3D Visuo-Tactile Representation Fusion
EchoVLA: Robotic Vision-Language-Action Model with Synergistic Declarative Memory for Mobile Manipulation
Multi-scale Mixture of World Models for Embodied Agents in Evolving Environments
Persistent Robot World Models: Stabilizing Multi-Step Rollouts via Reinforcement Learning
Tactile Modality Fusion for Vision-Language-Action Models
EgoTraj: Real-World Egocentric Human Trajectory
Personalization as Inverse Planning: Learning Latent Design Intents for Agentic Slide Generation via Structural Denoising
BeyondSight: Object Permanence for End-to-End Autonomous Driving
ExploreVLA: Dense World Modeling and Exploration for End-to-End Autonomous Driving
RL-AWB: Deep Reinforcement Learning for Auto White Balance Correction in Low-Light Night-time Scenes
Ranked Activation Shift for Post-hoc Out-of-Distribution Detection
Stabilizing Deep Reconstruction Operators with Contractive Anchoring
Estimating Individual Tree Height and Species from UAV Imagery
Any to Full: Prompting Depth Anything for Depth Completion in One Stage
Broadband Wide Field of View Imaging with Computational Mirrors
Physically Grounded Monocular Depth via Nanophotonic Wavefront Encoding
Video Generative Models as Geometry Learner
Parallax Portrait Matting
SynLF: Zero-Shot Metric Depth from Light Field Cameras via Physics-Grounded Synthesis
DP-BOA: Dirichlet-Process Birth-or-Assign for On-the-Fly Category Discovery
From Perspective to Fisheye Depth Estimation and Open-Vocabulary Segmentation
Leveraging Phase Information to Boost Unrolled Network Learning for Image Deblurring
EventSpecPS: Photometric Stereo with Multispectral Reflectance Using an Event Camera
Color Pass-Through via Camera-Display Coupling
A Mechanism-Driven Theory of Phase Transitions in Active Learning
Stable and Scalable Bundle Adjustment of Holistic 3D Structures
Cast and Attached Shadow Detection via Iterative Light and Geometry Reasoning
Geometry-Aware Visual Representation for Remaining Useful Life Prediction
Semantic Line Diffusion: Character-Consistent Line Art from text-annotated Storyboards
Don’t Mask Out the Background! Natural-Light Photometric Stereo via Illumination Reconstruction
mmIR: Frequency-Space Inverse Rendering for 3D Millimeter-Wave Radar ADC Synthesis
Pixel-wise Planarity for High-Precision Monocular Plane Segmentation
PixVOD: Pixel-Distributed Direct Visual Odometry and Depth Estimation
CrossFeat: Bridging Imaging Modalities in Feature Descriptor Space
Learning Ego-Centric BEV Representations from a Perspective-Privileged View: Cross-View Supervision for Online HD Map Construction
The 3D Mirage: Probing and Taming 3D Hallucinations
OREO: Fidelity Alignment in 3D Generation via On-the-fly Rendering-Editing Optimization
Real-Time LiDAR Gaussian Splatting SLAM via Geometry-Aware Covariance Coupling
Matryoshka Gaussian Splatting
Identity-Preserving Human Reconstruction from a Single Image via 3D Token Inference
StereoGS: Sparse-View 3D Gaussian Splatting via Stereo Priors
Global Pose Control for Generative View Synthesis in Normalized Object Coordinate Space
CARA: Collision-Aware Resolution Adaptation for Multiresolution Hash Encoding Based Image Fitting
SpectralSplats: Robust Differentiable Tracking via Spectral Moment Supervision
ViewSplat: View-Adaptive Dynamic Gaussian Splatting for Feed-Forward Synthesis
DreamWorld: Geometry-Grounded Video Diffusion for 3D-Consistent World Modeling
UniGeo: Unifying Geometric Constraints for Camera-Controllable Image Editing via Video Priors
MeGAS: Thermomechanical Dynamic Gaussian Splatting for Thermophysical Scene Editing
Reflection-aware generative novel view synthesis
DualDiff3D: Dual Structure-Appearance Diffusion Priors for Reliability-Enhanced 3D Gaussian Splatting
AutoWeather4D: Autonomous Driving Video Weather Conversion via G-Buffer Dual-Pass Editing
Render-FM: Feedforward Model for Real-time Photorealistic Volumetric Rendering
Rotate Your Character: Revisiting Video Diffusion Models for High-Quality 3D Character Generation
InstGS: Shared-Template Gaussian Instancing for Object-Redundancy-Free Rendering
AnchorSplat: Fast and Structure Consistent Detail Synthesis for Gaussian Splatting
AdaptiveSplat: Texture Aware Controllable 3D Gaussian Allocation for Feed-Forward Reconstruction
MetaView: Monocular Novel View Synthesis with Scale-Aware Implicit Geometry Priors
GlassGS: Geometry and Concept-Aware 3D Gaussian Splatting for Reflective Enclosures
ReconSplat: Generalizable 3D Scene Reconstruction Beyond Observed Views
LiteMatch: Lightweight Zero-Shot Stereo Matching via Cost Volume Stabilization
ReGen3D: Generalizable Unified Representation Learning for 3D Understanding
PointGT: Simultaneous Geometric and Textural Editing for Point-Based Representations
UniFusion: Sparse-View 4D Reconstruction via Unified Spatio-temporal Depth Alignment
WildSplat: Feedforward Gaussian Splatting from Unposed In-the-Wild Images
GaussianLens: Localized High-Resolution Reconstruction via On-Demand Gaussian Densification
LVSPM: Long Sequence View Synthesis and Pose Estimation Model
PDF-Omni: Poincaré Dual Disk Distortion Field-based Recurrent Update for Omnidirectional Stereo Matching
UniSim-SLAM: Feed-Forward SLAM with Unified Sim(3) Optimization
KISS-GS: 3D Gaussian Splatting Compression Kept Simple
SplatPainter: Interactive Authoring of 3D Gaussians from 2D Edits via Test-Time Training
Drop-In Perceptual Optimization for 3D Gaussian Splatting
SkipGS: Post-Densification Backward Skipping for Efficient 3DGS Training
F⁴Splat: Feed-Forward Predictive Densification for Feed-Forward 3D Gaussian Splatting
Targeted Structure Completion for Sparse-View 3D Reconstruction in Autonomous Driving
3D Gaussian Texture for Real-time Mesoscale Appearance Synthesis and Rendering
GeoV2V: Geometry-Grounded Video Diffusion Model for Driving Scene Generation
GlobalSplat: Efficient Feed-Forward 3D Gaussian Splatting via Global Scene Tokens
TriSplat: Adaptive Triplane for Sparse-View Large-Scale Scene Reconstruction
Improving Sparse-View 3DGS Generalization via Flat Minima Optimization
LEGO: Leveled Language Gaussian Splatting
Towards Alias-Free 4D Gaussian Representations with Motion-Aware Filtering
Novel View Synthesis as Video Completion
Wid3R: Wide Field-of-View 3D Reconstruction via Camera Model Conditioning
MLP Splatting: Object-Centric Neural Fields
Fourier Splatting: Generalized Fourier encoded primitives for scalable radiance fields
CubicSplat: Differentiable Vector Graphics via Error-Bounded Forward Relaxation
CAM3R: Camera-Agnostic Model for 3D Reconstruction
FixAnything: 3D-Consistent Rendering Refinement via Video Generative Priors
3D-ReGen: A Unified 3D Geometry Regeneration Framework
Extending a Large View Synthesis Model for Multi-view Panoptic Segmentation
EGGS: Explicitly Granular 3D Gaussian Splatting via Luma-Aware and Volume-Preserving Attribute Factorization
RayRoPE: Projective Ray Positional Encoding for Multi-view Attention
Recurrent Sinusoidal INRs for Efficient High-Fidelity Representation
X-SG2S: Safe and Generalizable Gaussian Splatting with X-dimensional Watermarks
RoomPlanner: Reachability-Aware View Sampling for Text-to-Room 3D Gaussian Splatting
MedGSSR: Generalizable Medical Image Super-Resolution 3D Reconstruction via Hierarchical Feed-forward Gaussian Splatting
Same Person, Different Depiction: Counterfactual Evaluation of Vision-Language Models on Individuals with Limb Deficiencies
Disentangling Pictorial Cue Understanding from Language Bias in VLMs via Depth Ordering Task
A4-Agent: An Agentic Framework for Zero-Shot Affordance Reasoning
GEM: Generative Supervision Helps Embodied Intelligence
CFM: Language-aligned Concept Foundation Model for Vision
SPOT-E: Test-Time Entropy Shaping with Visual Spotlights for Frozen VLMs
RAU: Reference-based Anatomical Understanding with Vision-Language Models
Trustworthy Image Authentication using Forensic Knowledge Graphs
Entropy-Gradient Grounding: Training-Free Evidence Retrieval in Vision-Language Models
AdaBoosting Text Prompts for Vision-Language Models
Bridging Visual Representation and Reinforcement Learning from Verifiable Rewards in Large Vision-Language Models
How Far Are Vision-Language Models from Constructing the Real World? A Benchmark for Physical Generative Reasoning
Mechanistic Finetuning of Vision-Language-Action Models via Few-Shot Demonstrations
Molmo-Point: Better Pointing for VLMs with Grounding Tokens
Zoom-IQA: Image Quality Assessment with Reliable Region-Aware Reasoning
Towards Reliable Medical Large Vision-Language Models via Counterfactual Preference Optimization
3D-Layout-R1: Structured Reasoning for Language-Instructed Spatial Editing
Visual Prompt Discovery via Semantic Exploration
CARE: Causally-Aligned Reasoning Exploration for Medical Large Language Models
Learning from Primitive: Probing Visual Reasoning of LVLMs via Counting
GAIA: A Data Flywheel System for Training GUI Test-Time Scaling Critic Models
Same Pool, Different Answer: Stable Best-of-N Selection for Vision-Language Models
Evaluating Reasoning Coherence in Video Generative Models with Text and Visual Hints
Restoring Linguistic Grounding in VLA Models via Train-Free Attention Recalibration
Dynamic Cluster Data Sampling for Efficient and Long-Tail-Aware Vision-Language Pre-training
On Test-Time Scaling for Vision-Language Models
SpatiO: Adaptive Test-Time Orchestration of Vision-Language Agents for Spatial Reasoning
StarDojo: Benchmarking Open-Ended Behaviors of Agentic Multimodal LLMs in Production–Living Simulations with Stardew Valley
BeTTER: Diagnose the Illusion of Embodied Reasoning in Vision-Language-Action Models
DEX-AR: A Dynamic Explainability Method for Autoregressive Vision-Language Models
URoPE: Universal Relative Position Embedding across Geometric Spaces
CulinaryCut: A Physics-aware Vision-Language-Action Benchmark for Food Cutting via Material Point Method
VERITAS: A Multi-agent Co-scientist for Verifiable Image-Derived Hypothesis Testing
From Gaze to Meaning: An AI Agent for Unified Zero-Shot Grounding and Explanation
Reinforcing Vision-Language Models for Image Quality Assessment with Grounding Process Rewards
MVPruner: Dynamic Token Pruning for Accelerating Multi-view Vision-Language Models in Autonomous Driving
Seeing Isn't Orienting: A Cognitively Grounded Hierarchical Benchmark for Object Orientation in MLLMs
Ceptor: Vision-Language Model-Infused Diverse Guidance for Detecting Anything
Gender Bias in Vision-Language In-Context Learning
Personalize Your Large Vision-language Models With In-context Prompt Tuning
LEGO-Puzzles: How Good Are MLLMs at Multi-Step Spatial Reasoning?
MMR-Bench: A Comprehensive Benchmark for Multimodal LLM Routing
ReShift: Aha-Moment-Driven Reasoning-Level Backdoor Attacks on Vision–Language Models
Open Your Eyes: Benchmarking the Detection of Fabricated Realities and Weaponized Ethics in VLMs
Evaluating and Understanding Model Editing for Medical Vision Language Models
CrossView: Can Vision-Language Models Reason Across Cameras?
VisReflect: Latent Visual Reflection for Fine-Grained Perception in Long Visual Context
V-REX: Benchmarking Exploratory Visual Reasoning via Chain-of-Questions
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding
VDC-Agent: When Video Detailed Captioners Evolve Themselves via Agentic Self-Reflection
Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction
Consistent Video-to-Video Translation via Explicit Correspondences
Multi-modal Knowledge Preserving Adapter for Embedding Backward Compatibility
Human-Level Accuracy, Non-Human Strategies: Revealing Model-Human Divergence in Video Physical Reasoning
Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs
OmniScript: Towards Audio-Visual Script Generation for Long-Form Cinematic Video
DART: Difficulty-Adaptive Routing for Zero-Shot Video Temporal Grounding
Cambrian-P: Pose-Grounded Video Understanding
End-to-End Training for Autoregressive Video Diffusion via Self-Resampling
Temporal and Cross-modal Alignment for Enhanced Audiovisual Video Captioning
EgoSAT: A Comprehensive Benchmark of Egocentric Streaming Interaction Understanding
OmniForcing: Unleashing Real-time Joint Audio-Visual Generation
Egocentric Procedure Parsing
StoryTeller: Training-Free Narrative Grounding for Long-Form Audio Description
ReQuest: Rethinking-based Question-Aware Frame Selection for Long-Form Video QA
HorizonRelight: Relighting Long-horizon Videos Consistently via Diffusion Transformers
Actor as Its Own Critic: Unifying Region Understanding and Localization via CycleGRPO
VisionCoach: Reinforcing Grounded Video Reasoning via Visual-Perception Prompting
EventSTU: Event-Guided Efficient Spatio-Temporal Understanding for Video-LLMs
Break Visual-Linguistic Asymmetry: Unleashing VLM's Cross-Modal Potential for General Face Forgery Detection
QCA: Query- and Content-Aware Keyframe Selection for Long Video Understanding
Don't Let the Video Speak: Audio-Contrastive Preference Optimization for Audio-visual Language Models
HyLaR: Hybrid Latent Reasoning with Decoupled Policy Optimization
Video-Oasis: Rethinking Evaluation of Video Understanding
Revisiting Weakly-Supervised Video Scene Graph Generation via Pair Affinity Learning
LongEgoRefer: A Benchmark for Long-Form Egocentric Video Referring Expression Comprehension
Learning from Failure: Inference-Time Self-Improvement for Computer-Use Agents
Towards Effective Long Video Understanding: Dynamic MAS Construction via Meta-Agent
Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning
PackForcing: Short Video Training Suffices for Long Video Sampling and Long Context Inference
VideoSearch-R1: Iterative Video Retrieval and Reasoning via Soft Query Refinement
Perceptual Projection Pruning: Diversity-Aware Video Token Pruning for Multimodal Large Language Models
RL from Physical Feedback: Aligning Large Motion Models with Humanoid Control
Less Data, Faster Convergence: Goal-Driven Data Optimization for Multimodal Instruction Tuning
SymbOmni: Evolving Agentic Omni Models via Symbolic Concept Learning
CurveStream: Boosting Streaming Video Understanding in MLLMs via Curvature-Aware Hierarchical Visual Memory Management
Less Tokens, Better Forecasts: Sparse Residual Routing for Efficient Weather Prediction
TRAM: Finetuning-Free Test-Time Adaptation for Generalized Face Anti-Spoofing with Only a Few Bonafide Samples
MegaFlow: Zero-Shot Large Displacement Optical Flow
From Local Windows to Adaptive Candidates via Individualized Exploratory: Rethinking Attention for Image Super-Resolution
Dual-Prior Guided Null-Space Learning with Mixture-of-Splines for Arbitrary Medical Slice Super-Resolution
Scaling Whole-Slide Pathology Foundation Model Pretraining with Billions Off-the-Shelf Tokens
FreqPhys: Repurposing Implicit Physiological Frequency Prior for Robust Remote Photoplethysmography
STARLINC: Satellite Trail Artifact Removal using Inter-Frame Correlation
LIIFusion: Coarse-to-fine Framework for Generative MEF via Implicit Neural Representation
Ice Cloud Geometry Retrieval with Calibrated Uncertainty from Passive Satellite Imagery
CLARITY: Medical World Model for Guiding Treatment Decisions by Simulating Context-Aware Disease Trajectories in Latent Space
LineGraph2Road: Structural Graph Reasoning on Line Graphs for Road Network Extraction
WARP: Wide Attention with Rich Projections for Image Super-Resolution
MorphJEPA: Morphology-Aware Latent Prediction for Hyperspectral Images
MOOZY: A Patient-First Foundation Model for Computational Pathology
Zero-shot Depth from Defocus
HybridSim: A Physics–Learning Hybrid Digital Twin for mmWave Human Sensing
HSFM: Hard-Set-Guided Feature-Space Meta-Learning for Robust Classification under Spurious Correlations
Task-driven Processing with Coarse-to-Fine Glimpse-based Active Perception
EchoSonar-R: A Multi-View Reasoning-Enabled Model for Disease Classification and Report Generation in Echocardiography
Integrated Forward–Inverse Network for Reconstruction for Lensless Image Reconstruction
RobustRDP: Advancing Reaction Diagram Parsing via Synthetic-to-Real Data Scaling and Robustness-Oriented Training
MedSPOT: A Workflow-Aware Sequential Grounding Benchmark for Clinical GUI
Linguistically-Aligned and Visually-Grounded Preference Optimization for Clinically-Augmented Medical Report Generation
Off the Planckian Locus: Using 2D Chromaticity to Improve In-Camera Color
OrthoEraser: Coupled-Neuron Orthogonal Projection for Concept Erasure
On the Reliability of Cue Conflict and Beyond
This Looks Distinctly Like That: Grounding Interpretable Recognition in Stiefel Geometry against Neural Collapse
Data Circuit Breaker: Identifying Training, Test, and Generated Data in Image Generative Models
Denoised Variance-Based Pruning with Optimal Brain Bias Compensation
Compact and Structurally Transparent Cervical Cytology with Geometry-Driven Features and Closed-Form Attention
Different Changes Require Different Reasoning: Change-Type-Specialized Experts for Robust Change Captioning
Are GUI Agents Focused Enough? Automated Distraction via Semantic-level UI Element Injection
Frequency Director: Learnable Mixture of Frequency Experts for Unified Concealed Scene Segmentation
Intrinsically Stable Spiking Neural Networks: Overcoming the Performance Barrier in the Absence of Batch Normalization
R-ESC: Robustly Erasing Space Concepts via Stochastic Feature Remapping
Defending from GeoLocalization through Adversarial Road Trips
Weight Feedback Computes the Exact Jacobian Transpose in Modern Deep Networks
Geometry-Anchored Transport Framework for Exemplar-Free Class-Incremental Learning
Geometry Aware Reliable Instance Selection for Noisy Partial Label Learning
ORBIT: Overcoming Hallucination Risks via Bi-manifold Interaction and Traction
Enhancing Pretrained Model-based Continual Representation Learning via Guided Random Projection
BackTranslation2.0 - A Linguistically Motivated Metric to Assess Sign Language Production
AracNet: Revealing Debiasing Signals across Layers with Shallow Monitors
Exposing Implicit Vulnerabilities in Text-to-Image Models via Adversarial Agentic Probing
Closing the Capacity–Convergence Gap: Globally Optimal Configuration of Implicit Neural Representations
RoME: Robust Mixture of Low-Rank Experts against Multiple Adversarial Perturbations
Structured-Noise Masked Modeling for Video, Audio and Beyond
Causal Intervention in Concept Bottleneck Models
Leveraging Cross-Modal Knowledge Transfer for Knowledge-Aware Concept Customization
Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models
Diffusion-Based Immersive Visual Reasoning
Look Less, Think Faster: Joint Token-Compute Adaptation for Multimodal LLMs
Enlightening Photographic Style Transfer with a Self-Supervised Photographic Embedding
Verifying Cancer Segmentation in Vision Transformers via Internal Concepts
Rethinking Attention Reallocation for Multimodal Emotion Recognition
Dynamic Image Prompt Adapter for Scalable Zero-shot Personalized Text-to-Image Generation
Invisible Shortcuts: Why Vision Encoders Know Your Camera
World Knowledge in the Weights: Reading Concept Circuits of Vision Transformers
SpatialBoost: Enhancing Visual Representation through Language-Guided Reasoning
CL4D: Contrastive Language–4D Pretraining for Vision-Language Reasoning in Dynamic Scenes
From Drop-off to Recovery: A Mechanistic Analysis of Segmentation in MLLMs
Why Can Accurate Models Be Learned from Inaccurate Annotations?
MIRROR: Aligning Semantic Relations from Language to Image via Gromov--Wasserstein
VIVAS: Vitalizing Visual Perception in VLM Pre-training via Vision-language Unified Autoregressive Supervision
Background Blurring Matters: Improving Visual Grounding by Merging Text-Irrelevant Tokens
Contrastive-Guided Self-Supervised Latent Visual Reasoning for Hallucination Mitigation
What CLIP Knows but Cannot Say: Recovering Negation from Frozen Intermediate Features
Tesselating The Earth
Seeing Through Circuits: Faithful Mechanistic Interpretability for Vision Transformers
HyFL-CLIP: Hyperbolic Fine-Tuning of CLIP for Robust Long-Context Understanding
Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens
AFFMAE: Scalable Vision Pre-Training for High-Resolution Microscopy Segmentation on Desktop Hardware
When Token Compression Breaks: Structural Pruning vs. Token Reduction for Robust ViT Segmentation under High Compression
ECC: Encoder-Centric Corruption for Fine-Grained Vision in VLMs
Practice Makes Perfect: From Explicit Decomposition to Reinforced Latent Planning in Text-to-Human Motion
What VGGT Knows About Overlap: Probing Geometric Foundation Models for Co-Visibility
Make Geometry Matter for Spatial Reasoning
Probing the 3D Object-Level Understanding of Pre-Trained Detection Transformers
TETO: Tracking Events with Teacher Observation for Motion Estimation and Frame Interpolation
StreamTalk: Streaming Co-Speech Gesture Generation with Key-Pose Anchoring
Syn4D: A Multiview Synthetic 4D Dataset
PointLAM: Local Attentive Mamba for Efficient Point-based 3D Object Detection
MemPose: Category-level Object Pose Estimation with Memory
RoMa v2: Harder Better Faster Denser Feature Matching
TACO-Net: Topological Signatures Triumph in 3D Object Classification
360Anything: Geometry-Free Lifting of Images and Videos to 360°
Rolling Shutter Relative Pose Estimation Made Practical
Reconstructing Humans and Objects in Interaction using Large Reconstruction Models
General Self-Calibration with Varying Intrinsics
Pointer-CAD v2: Plan-Then-Construct CAD Generation with Dimension-Aware Parametric Precision
UniFlow: Zero-Shot LiDAR Scene Flow for Autonomous Driving
Delaunay Canopy: Building Wireframe Reconstruction from Airborne LiDAR Point Clouds via Delaunay Graph
Ray-Path-Aware Virtual Point Removal on 2D Layer-Wise Nearest Point Map
FreeFlow: A Bias-free Hierarchical Transformer for Optical Flow Estimation
ObjectForesight: Predicting 3D Object Trajectories from Human Videos
Sequential Visual Place Recognition: Exploiting Trajectory Priors for Robust Localization
Towards Reconfigurable Visual Feature Compression
PriorPose: Reference-Guided Joint Deformation and Alignment for Category-Level Object Pose Estimation
Learning Global Camera Poses from Noisy View-Graphs for Structure from Motion
OpenCVL: An Open, Diverse, and Large-Scale Dataset for Fine-Grained Cross-View Localization
Social-Mamba: Socially-Aware Trajectory Forecasting with State-Space Models
E-MOTION: A Dataset for Event-Based Scene Flow Estimation with Independent Moving Objects
Estimating Velocity and Spin of Spherical Objects from Rolling-Shutter Image(s)
MMEarth-Bench: Global Model Adaptation via Multimodal Test-Time Training
Dense Dynamic Scene Reconstruction and Camera Pose Estimation from Multi-View Videos
OmniDS: Dual-Stream Context Fusion for Omnidirectional Depth from Fisheye Cameras
Training-free Controllable Motion Generation under Heterogeneous Constraints
RADmesh: Remesh-Aware Mesh Deformation
Sparsity-Inducing Divergence Losses for Biometric Verification
MoBa-GS: Learning a Spatially-Varying Motion Basis over a Dynamic Canonical Space for 4D Reconstruction
Tempo-SAM3D: Monocular Video to 4D via Temporal Memory-Guided Generation
Analytic Bayesian Uncertainty for LiDAR Segmentation: A Single-pass Generative Approach
InSpace: Structure-Aware 3D Indoor Scene Generation from a Single 360° Image
NEOMAP: Novel-View Synthesis via Noise Initialization by Manifold Alternating Projection
LibraGen: Playing a Balance Game in Subject-Driven Video Generation
Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length
GO-Renderer: Generative Object Rendering with 3D-aware Controllable Video Diffusion Models
Walking in the Implicit: Interactive World Exploration via Neural Scene Representation
MonoArt: Progressive Structural Reasoning for Monocular Articulated 3D Reconstruction
One Video, One World: Turning Monocular Video into Physical 4D Scenes
DualCamCtrl: Dual-Branch Diffusion Model for Geometry-Aware Camera-Controlled Video Generation
DiffProxy: Multi-View Human Mesh Recovery via Diffusion-Generated Dense Proxies
Human Mesh Modeling for Anny Body
TryOnCrafter: Unleashing Camera Trajectories for Realistic Video Virtual Try-on via a Renderable 4D Try-on Proxy
Predictive Structure Improves Video Diffusion Dynamics
CoSyncDiT: Cognitive Synchronous Diffusion Transformer for Movie Dubbing
What Moves? Localized Motion Representations for Compositional Scene Control
Video2Reaction: Mapping Video to Audience Reaction Distribution in the Wild
MetaPoint: Unlocking Precise Spatial Control in Visual Generation
Measuring 3D Spatial Geometric Consistency in Dynamic Video Generation
Meric: A Unified Framework for Multimodal Music Generation and Retrieval via Representation Space Anchoring
Scalable Cross-embodiment Dexterous Grasping via Morphology-Prior Diffusion
DiffHDR: Re-Exposing LDR Videos with Video Diffusion Models
InfiniteDance: Scalable 3D Dance Generation Towards in-the-wild Generalization
Dynamic World Generation Made Efficient
Semantic-Aware, Physics-Informed, Geometry-Grounded Weather Synthesis
ID-PreFeR: ID-Preserving Face Restoration with Mixed Data Quality
3D FaceShell: Attribute Transfer in 3D Face Avatars as a VLM Defense Mechanism
OmniDance: Multimodal Driven Dance Video Generation with Large-scale Internet Data
OctWorld: Long-Range World-Consistent Video Generation with Octree-based 3D Mapping
Taming Camera-Controlled Video Generation with Verifiable Geometry Reward
SMG: Semantic Motion Graph for Monocular Dynamic Gaussian Splatting
MemLearner: Learning to Query Context Memory for Video World Models
FrozenDrive: Zero-Shot Text-Guided Driving Scene Generation and Data Augmentation with Parameter-Free Frozen Diffusion Model
ICDepth: Taming Video Diffusion Models for Video Depth Estimation via In-Context Conditioning
AerialMetric: Benchmarking and Adapting UAV Monocular Metric Depth Estimation in the Real World
ReCamDriving: LiDAR-Free Camera-Controlled Video Synthesis for Novel Trajectories
PIAvatar: Physically Interactive Avatars via Deformation Gradient Decoupling
InsertAnywhere: Geometrically Grounded and Optics-Aware Video Object Insertion
ActionParty: Multi-Subject Action Binding in Generative Video Games
VERTIGO: Visual Preference Optimization for Cinematic Camera Generation
Generative Action Tell-Tales: Assessing Human Motion in Synthesized Videos
CortexVideo: A Semantic-Spatial Dual-Anchor Framework for High-Fidelity fMRI-to-Video Reconstruction
OmniHuman: A Large-scale Dataset and Benchmark for Human-Centric Video Generation
Head Avatars with Dynamic Explicit Hair
DCARL: A Divide-and-Conquer Framework for Autoregressive Long-Trajectory Video Generation
FlexAM: Flexible Appearance-Motion Decomposition for Versatile Video Generation Control
Monocular Models are Strong Learners for Multi-View Human Mesh Recovery
OmniColor: A Unified Framework for Multi-modal Lineart Colorization
Tuning-free Visual Effect Transfer across Videos
ETCH-X: Robustify Expressive Body Fitting to Clothed Humans with Composable Synthetic Data
SIMSplat: Language-Aligned 4D Gaussian Splatting for Driving Scenario Generation
Reconstruction by Generation: 3D Multi-Object Scene Reconstruction from Sparse Observations
The Dynamic Prior: Understanding 3D Structures for Casual Dynamic Videos
PhysConvex: Physics-Informed Dynamic Convex Fields for Reconstruction and Simulation
NeuIDO: Neural Intrinsic Dynamics Operator for Physics-Informed 4D World Models
BeyondMasks: Evaluating Causal and Physical Consistency in Video Object Removal
AutoPhyX: Automatic Text-Condition Physics Property Generation
Who Does What and Where to Go: Orthogonal Alignment and Hierarchical Planning for Multi-Entity Trajectories
MorphGS: Morphology-Adaptive Articulated Motion Transfer from Videos
ROSE: Real-Time Open-World Scene Understanding from Monocular Video via Compact Multimodal 4D Scene Graphs
Interaction-Aware 4D Gaussian Splatting for Dynamic Hand-Object Interaction Reconstruction
SA-V2V: Training-Free Subject-Aware Video-to-Video Personalization
D-Rex : Diffusion Rendering for Relightable Expressive Avatars
LUNA: Learning Universal 3D Human Animation Beyond Skinning
Towards Flexible, Natural, Efficient Interaction for Conversational Talking Face Generation
SICAGE: Speaker-Independent Culture-Aware Gesture Generation using TED4C-L Dataset
GraphVid: Interactive Graph-Controllable Video Generation
SignRefine: Adapting Foundational Video Models for Sign Language Generation
JointHOI: Jointly Generating Contact Maps Enhances Hand Object Interaction Generation
RealDyadic: Synthesizing Realistic Dyadic 3D Dialogue with Neural Appearance Priors
Taming Dynamic Clutter: Variance-Driven Adaptive Gain Control for Bio-inspired Small Target Detection
Learning Implicit Constitutive Laws for Dynamic 3D Gaussian Splatting from Monocular Videos
FlexComposer: Unified Video Compositing from Images to Dynamic Footage with Flexible Trajectory Control
WristMimic: Full-Body Humanoid Control with Wrist-Guided Manipulation
SplatCtrlA: Generalizable Single Image to Fully Controllable 3D Avatar
Beyond Pixel Mimicry: Disentangled Self-Similarity Rewards for Diverse Subject-Driven Generation
From Synchrony to Sequence: Exo-to-Ego Generation via Interpolation
RayDer: Scalable Self-Supervised Novel View Synthesis from Real-World Video
Cycle-World: Mitigating Error Accumulation in Long-term Video World Models via Reverse-Prediction Cycle Consistency
LooseControlVideo: Directorial Video Control using Spatial Blocking
DynEval: Holistic Evaluations of T2I Generative Models in the Wild
Physics Question Scene Graph: Fine-grained Evaluation of Physical Plausibility in Text-to-Video Generation
TAQ: Static-Deployable Temporal-Aware Quantization for Real-World Video Super-Resolution
DIVER: Disentangling Camera–Object and Active–Passive Motion for Video Generation
A Physics-Grounded Benchmark for Multi-Agent Dynamics in World Models
ExpertEdit: Learning Skill-Aware Motion Editing from Expert Videos
Music-to-Dance Generation via Atomic Movements
MSEditor: Toward Consistent Multi-Shot Video Editing
YeTI: You Only Need Two Noisy Images for Real-World sRGB Noise Generation
TextFace: Compositional Text-Guided Identity Preserving Face Synthesis for Face Recognition
IRIS: A Real-World Benchmark for Inverse Recovery and Identification of Physical Dynamic Systems from Monocular Video
FlowLess: Controlling Abstract Image Generation
WorldCache: Content-Aware Caching for Accelerated Video World Models
Video Generation Models are General-Purpose Vision Learners
Scaling Dense Prediction with Latent Decoding
Kinematics-Agnostic 3D Human Motion Prediction via Equivariant Latent Diffusion
Complex-Valued 2D Gaussian Representation for Computer-Generated Holography
Logit Refiner: Improving Visual Autoregressive Models via Intra-Scale Dependency Modeling
On the Diffusibility of High-Dimensional Latents
Diffusion-Based Material Regularization for Physics-Based Inverse Rendering
ReDesign: Recovering Editable Design Structures from Raster Images via Agentic Decomposition
Heat Kernel Textures -- the Geodesic Gaussians That Do Not Splat
ShellMaker: Language-Guided Exterior Completion under Structural Constraints
EVAR: Edge Visual Autoregressive Models via Principled Pruning
FlowFace: Rectifying Identity Conditioning with Riemannian Geometry for Face Generation
Bridging Online and Offline Handwriting via Differentiable Physical Rendering
Exo2EgoPolicy: Pose-Aligned Cross-View Policy Learning
EvoTok: A Unified Image Tokenizer via Residual Latent Evolution for Visual Understanding and Generation
GenSP: Consistent Spherical Parameterization via Learning Shape Generative Models
Multi-View Foundation Models
DreamCAD: Scaling Multi-modal CAD Generation using Differentiable Parametric Surfaces
APT: Anchor-aligned Perturbations for Tamper Localization in Fully Regenerated Images
SimFlow: Simplified and End-to-End Training of Latent Normalizing Flows
TopoFuse: Topology-Aware Tri-Planar Fusion for 3D Cryo-Electron Tomography Segmentation
Text-based Tactile Graphics Generation for the Visually Impaired
Pixel Ignores, Superpixel Sees: Adverse Weather Image Restoration via Semantic-Center SSM
Repurposing Geometric Foundation Models for Multi-view Diffusion
Stitched Embeddings: A Unified Latent Space for 3D Garments and 2D Patterns
MoScale: Autoregressive Next-Scale Prediction for Human Motion Generation and Editing
Query-Kontext: An Unified Multimodal Model for Image Generation and Editing
Next-Frame Decoding for Ultra-Low-Bitrate Image Compression with Video Diffusion Priors
Text-Conditioned Background Generation for Editable Multi-Layer Documents
Learning to Stylize by Learning to Destylize: A Scalable Paradigm for Supervised Style Transfer
NoiseTilt: Noise-Tilted Reverse Kernels for Diffusion Reward Alignment
MPO: Single-Stream Policy Optimization for Efficient Text-to-Image Alignment
AnyFlow: Any-Step Video Diffusion Model with On-Policy Flow Map Distillation
Cheers: Decoupling Patch Details from Semantic Representations Enables Unified Multimodal Comprehension and Generation
From Draft to Draft-Free: One-Step Video Object Removal via Privileged Distillation and Fast Planting
LangDriveCTRL: Natural Language Controllable Driving Scene Editing with Multi-modal Agents
WeEdit: A Dataset, Benchmark and Glyph-Guided Framework for Text-centric Image Editing
UniReflect: Self-Reflection Tuning for Unified Multimodal Understanding and Generation
Rethinking Reward Signals in Video GRPO: When Scores Become Targets
DisRM: Reward Modeling as Discriminative Prediction
Finite Difference Flow Optimization for RL Post-Training of Text-to-Image Models
Rethinking Visual Privacy: A Compositional Privacy Risk Framework for Severity Assessment with VLMs
Consistent Feature Transport for Image Relighting
SupGRPO: Enhancing GRPO with Matching-based Online SFT for Text Spotting
i-Design: Step-by-Step Graphic Layout Design with Progressive Aesthetic Policy Optimization
InstantRetouch: Personalized Image Retouching without Test-time Fine-tuning
Few-Shot Synthetic Image Attribution: Identifying Unseen Generators with Limited Samples
Unified Removal of Raindrops and Reflections: A New Benchmark and A Novel Pipeline
Bottom-up modeling of repeated elements via single image analysis-by-synthesis
SceneDiff: A Benchmark and Method for Multiview Object Change Detection
Robust Self-Supervised Cross-Modal Super-Resolution against Real-World Misaligned Observations
Making Partial-Label Datasets Easier: A Simple Yet Highly Effective Data Augmentation for Deep Partial-Label Learning
CLUE-VAD: Structured Semantic Clues for Understanding Explainable Events in Video Anomaly Detection
GH-ESD: Grounded Hypothesis-Driven Error Slice Discovery for Instance-Level Vision Tasks
Online 3D Instance Segmentation at task-oriented granularity with Unposed Monocular Video
CAR-MIL: Counterfactual Attention Regularization for Multiple Instance Learning
Pseudo-Stereo Inputs: A Solution to the Occlusion Challenge in Self-Supervised Stereo Matching
Rapidly Deploying On-Device Eye Tracking by Distilling Visual Foundation Models
DETRPose: Real-Time End-to-End Multi-Person Pose Estimation via Modified Transformer Decoder and Novel Denoising Keypoints
Diffusion-based dual-view reflection removal
ParaFlow: Parallel Sampling for Flow Matching Models
Direct Autoregressive Diffusion Distillation via Error-aware Causal Pretraining
Training-Free Refinement of Flow Matching with Divergence-based Sampling
Discrete Diffusion Bridges for Spatiotemporally Aligned Image Translation and Generation
AViTS:Adaptive Spatiotemporal Token Selection for Efficient Dynamic-Resolution Generation
Expert Weaving: Marrying Masked AutoRegressive and Diffusion Models for Unified Image Restoration
V-Co: A Closer Look at Visual Representation Alignment via Co-Denoising
Efficient and High-Quality Depth Estimation via Pixel-Space Diffusion with Linear Attention
Prompt2Effect: Training-Free LoRA Synthesis for Controllable Video Effects
Allo{SR}2: Rectifying One-Step Super-Resolution to Stay Real via Allomorphic Generative Flows
UltraGen: Efficient Ultra-High-Resolution Image Generation with Hierarchical Local Attention
Ctrl-Z Sampling: Scaling Diffusion Sampling with Controlled Random Zigzag Explorations
UNITY: Attention Flow Networks for Adaptive Conditioning in Diffusion
D3F-IR: Dual-Domain Deterministic Flow Matching for Visible-to-Infrared Translation
AccelAes: Accelerating Diffusion Transformers for Training-Free Aesthetic-Enhanced Image Generation
Jumping the Landing Phase: Noise Variance Matching Enables Accurate Few-Step Inversion
FreqOrtho-SR: Frequency-Guided Orthogonal Expert Learning for Real-World Image Super-Resolution
GMODiff: One-Step Gain Map Refinement with Diffusion Priors for Efficient HDR Reconstruction
Beyond Prompts: Unconditional 3D Inversion for Out-of-Distribution Shapes
TanGO: Training-Free 3D Editing via Tangent-Space Guidance and Optimization
Unified Backbone Refinement for Diffusion Models via Internal-Latent Analysis
LGD-Net: Leader-Guided Cross-Modal Dynamics for Hyperspectral and Panchromatic Image Fusion
Why Can't I Open My Drawer? Mitigating Object-Driven Shortcuts in Zero-Shot Compositional Action Recognition
Online Versatile Incremental Learning: Towards Class and Domain-Agnostic Adaptation at Any Time
CHIMERA: Adaptive Cache Injection and Semantic Anchor Prompting for Zero-shot Image Morphing with Morphing-oriented Metrics
Enhancing prompt-image alignment evaluations via cyclic mutual information maximization
Learning Structured Visual Compositional Representations for Weakly Supervised Referring Expression Comprehension
Dual-Margin Embedding for Fine-Grained Long-Tailed Plant Taxonomy
Context-Aware Joint Alignment for Cross-Scene Hyperspectral Image Classification
DGSeg: Dynamic Gating of Semantic-Spatial Guided Predictions for Reasoning Segmentation
What is the Right Embedding Space for Contrastive Learning in Referring Expression Counting?
Rethinking Prototype-based Similarity Learning for Few-Shot Object Detection
CROSS: Cascaded Distillation and Dual-Constraint Grounding for Remote Sensing Referring Segmentation
ReynoldsFlow: Physics-Inspired Spatiotemporal Flow Representation for Video Understanding
SPLIT: Training-Free AI-Generated and Partially Edited Video Detection via Spatial Patch‑Level Incoherence and Temporal Roughness
FAIR: Feature-Augmented Implicit Regularization for AI-generated Fake Image Detection
OPAL: Orthonormal Prototype Alignment Learning for Interpretable Image Classification
Can Vision Models Truly Forget? Mirage: Representation-Level Certification of Visual Unlearning
Beyond Disjoint Tasks: Towards More Natural Continual Learning for Vision-Language Models
Shared LoRA Subspaces for almost Strict Continual Learning
VLOD-TTA: Test-Time Adaptation of Vision-Language Object Detectors
MMLoP: Multi-Modal Low-Rank Prompting for Efficient Vision-Language Adaptation
TouchAnything: Diffusion-Guided 3D Reconstruction from Sparse Robot Touches
FILT3R: Latent State Adaptive Kalman Filter for Streaming 3D Reconstruction
White Aggregation and Restoration for Few-shot 3D Point Cloud Semantic Segmentation
Pano3D: Unified 3D Reconstruction and Panoptic Segmentation
SHReg: Strictly Rotation-Equivariant Point Cloud Registration via Spherical Harmonics
Towards Practical Lossless Neural Compression for LiDAR Point Clouds
BitRIC: Efficient Neural Compression of LiDAR Range Images via Hierarchical Bitplanes
MARché: Fast Masked Autoregressive Image Generation with Cache-Aware Attention
Bootstrapping Articulated 3D Reconstruction from 2D Image Collections
DRS-VPT: Directly Re-localizing in Scenes using a Vision and Point Transformer
Seeing Through the Weights: Privacy Leakage in Scene Coordinate Regression
Multiple Images Distract Large Multimodal Models via Attention Fragmentation
Evidence-Backed Video Question Answering
Learning Active Perception for Pixel-Space Reasoning via Visual-Intent Stratified GRPO
ViTexQA: A Multi-Frame Temporal Perception Dataset for Video Text Question Answering
Starve to Perceive: Taming Lazy Perception in VLMs with Constrained Visual Bandwidth
Experts-Guided Unbalanced Optimal Transport for ISP Learning from Unpaired and/or Paired Data
HumanOmni-Speaker: Identifying Who said What and When
Enhancing Alignment for Unified Multimodal Models via Semantically-Grounded Supervision
VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning
SEERBench: A Spatial Ego-Exo Reasoning Benchmark for MLLMs with a Simple Yet Effective Baseline
UniRec-0.1B: Unified Text and Formula Recognition with 0.1B Parameters
Bridging Vision and Language Concepts through Optimal Transport Semantic Flow
InstrAct: Towards Action-Centric Understanding in Instructional Videos
CoverPrune: Coverage-Driven Token Pruning for 3D VLMs via Optimal Transport
Efficient Document Tampering Localization with Multi-Level Discrepancy Features and Unified DCT–Quantization Embedding
ELVA: Exploring Ranking-Driven Universal Multimodal Retrieval
Sink-Token-Aware Pruning for Fine-Grained Video Understanding in Efficient Video LLMs
When Sinks Help or Hurt: Unified Framework for Attention Sink in MLLMs
The Telephone Game: Evaluating Semantic Drift in Unified Models
SE-DETR: Explicit Semantic Exploration for Generalizability and Distinguishability in Video Temporal Grounding
How to Teach Large Multimodal Models New Skills
Identifying and Resolving Pitfalls of Knowledge-Based VQA Benchmarks: Auditing, Repairing, and Augmenting
Event-Driven Video Generation
Beyond Dense Futures: World Models as Structured Planners for Robotic Manipulation
Plug-and-Play Traffic Element Awareness for End-to-End Autonomous Driving
Spatial Amsan: A Benchmark for Perception-Grounded Spatial Reasoning and Action Evaluation in Egocentric Manipulation
Dual-Anchoring: Addressing State Drift in Vision-Language Navigation
MOJITO: Modal Joint Learning for Unified End-to-End Autonomous Driving
ESTANet: Efficient Online Error Detection in Procedural Videos via Prediction Inconsistency
Understanding Cross-Rig Generalization in Automotive Perception: a Multi-Rig Benchmark and Rig Variation Metrics
Unordered Landmark Visual Navigation
R3DP: Real-Time 3D-Aware Policy for Embodied Manipulation
Humanoid Whole-Body Manipulation via Active Spatial Brain and Generalizable Action Cerebellum
Physically Grounded 3D Generative Reconstruction under Hand Occlusion using Proprioception and Multi-Contact Touch
Robobench: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models as Embodied Brain
Cooking beyond Frames: A Stereo Event Camera Dataset in the Kitchen
Boxer: Robust Lifting of Open-World 2D Bounding Boxes to 3D
DVG-WM: Disentangled Video Generation Enables Efficient Embodied World Model for Robotic Manipulation
Rule-VLN: Bridging Perception and Compliance via Semantic Reasoning and Geometric Rectification
Agent-OBJ: Prompt-Driven 3D Adversaries for Multi-Modal Perception
GameWorlds: Towards Standardized and Verifiable Evaluation of Multimodal Game Agents
HSDF-Lane: Height-Aligned Signed Distance Field with Semantic Lane Prior for 3D Lane Detection
AeroVLA: A Vision-Language-Action Model for UAV Navigation via Minimalist End-to-End Control
KineBench: Benchmarking Embodied World Models via IDM-Free Kinematic Grounding
ZAP: Zero-Shot Assembly Planning with Large Language Models
VIPS: Vehicle-Infrastructure Cooperative Planning Benchmark via Pseudo-Simulation
Beyond Inpainting: Unleash 3D Understanding for Stable Camera-Controlled Video Re-rendering
Anchored Video Generation: Decoupling Scene Construction and Temporal Synthesis in Text-to-Video Diffusion Models
Look But Don't Touch with Sparse Autoencoders for Unlearning in Diffusion Models
MotionEditGS: Editing Motion and Appearance of 4D Scenes from Monocular Video via Semantically Anchored Gaussians
Show Me Examples: Inferring Visual Concepts from Image Sets
Large-Scale Light Field Synthesis from Videos Enables Geometrically Consistent Bokeh Editing
Keep Your Friends Close, and the Right Neighbours Closer: Disaster-Conditioned Kernel-Regularized Graph Attention for Building Damage Classification
P-CORE: Self-Supervised Surface Consistency for Point-Based Neural Editing
Beyond Isolated Scans: Cross-Phase Alignment of Structure and Topology for 3D Medical Pretraining
PrintAnything: Learning Geometric Plan Map for 3D Printing G-code Generation from Unoriented Point Clouds
H-SFP: Hierarchical Federated Learning with Decoupled Split-Model Prototyping
Flash-Refine: Frustum-Guided Local Incremental Learning for Efficient 3D Gaussian Splatting Completion
MLVC: A Multi-platform Learned Video Codec for Real-World Deployment
We use cookies to store which papers have been visited.
I agree
Successful Page Load
ECCV uses cookies for essential functions only. We do not sell your personal information.
Our Privacy Policy »
Accept
We use cookies to store which papers have been visited.
I agree