| 201 |
SAGE: Synchronized Action-Gaze Recognition and Anticipation for Human Behavior Understanding
Chenyi Kuang, Nakul Agarwal
|
|
cs.CV
|
0 |
27 days ago |
| 202 |
InSpace: Structure-Aware 3D Indoor Scene Generation from a Single 360° Image
Gwanhyeong Koo, Hyunsu Kim, ... (+6 more)
|
|
cs.CV
|
0 |
27 days ago |
| 203 |
Reward Lightning: Fast Video Generation via Homologous Preference Distillation
Jiaxiang Cheng, Bing Ma, ... (+5 more)
|
|
cs.CV
|
0 |
27 days ago |
| 204 |
DICT: Data Injection and Contrastive Trajectory Refinement for Conditional Image Generation with Diffusion Models
Chunnan Shang, Xin Zhang, ... (+2 more)
|
|
cs.CV
|
0 |
27 days ago |
| 205 |
BAT3R: Bootstrapping Articulated 3D Reconstruction from 2D Image Collections
Jakub Zadrozny, Oisin Mac Aodha, Hakan Bilen
|
|
cs.CV
|
0 |
27 days ago |
| 206 |
Q-TriM: Question-Guided Tri-Modal Attention for Audio-Visual Question Answering
SungHun Kim, SeungJun Baek
|
|
cs.CV
|
0 |
27 days ago |
| 207 |
FDR-Occ: Factorized Dense Routing for Full-Spectrum 3D Occupancy Prediction
Dubing Chen, Huan Zheng, ... (+6 more)
|
|
cs.CV
|
0 |
27 days ago |
| 208 |
Global Logic and Local Search: Dual-Stream Multimodal In-Context Learning for Verifiable Industrial Anomaly Detection
Runzhi Deng, Yundi Hu, ... (+6 more)
|
|
cs.CV
|
0 |
27 days ago |
| 209 |
City-Level 3D Surface Reconstruction with Viewpoint Orientation Partitioning and Scene Completion
Liang Han, Wenyuan Zhang, ... (+3 more)
|
|
cs.CV
|
0 |
27 days ago |
| 210 |
Self-Improving Diffusion Classifiers with Minority Preference Optimization
Hyunsoo Kim, Jungmyung Wi, ... (+3 more)
|
|
cs.CV
|
0 |
27 days ago |
| 211 |
Sparse-View Surface Reconstruction using Gaussian Splatting through High-Confidence Depth Propagation with Normal Priors
Liang Han, Bangcai Wei, ... (+3 more)
|
|
cs.CV
|
0 |
27 days ago |
| 212 |
Moonstone: A Multimodal Foundation Model and Benchmark for Lunar Remote Sensing
Ayush Prasad, Swarnalee Mazumder
|
|
cs.CV
|
0 |
28 days ago |
| 213 |
SAF3R: Dynamic Sparse Attention for Feed-Forward 3D Reconstruction Transformers
Jianing Deng, Yuanzhe Li, ... (+5 more)
|
|
cs.CV
|
0 |
28 days ago |
| 214 |
Token-Based Affordance Grounding with Large Vision-Language Models
Seung Il Lee, Qinqian Lei, ... (+5 more)
|
|
cs.CV
|
0 |
28 days ago |
| 215 |
WorldBagel: Uncovering the Power of Unified Multimodal Models for Vision-Language-Action-World Modeling
Zelin Zhao, Min Shi, ... (+6 more)
|
|
cs.CV
|
0 |
28 days ago |
| 216 |
GrowFields: Compositional 4D Neural Fields for Topology-Changing Plant Growth
Joaquin Gajardo, Michele Volpi, ... (+4 more)
|
|
cs.CV
|
0 |
28 days ago |
| 217 |
Defending from GeoLocalization through Adversarial Road Trips
Niccolò Niccoli, Federico Becattini, Lorenzo Seidenari
|
|
cs.CV
|
0 |
28 days ago |
| 218 |
ExpoMotion: A Large-Scale Benchmark and A Householder Projection Network for Multi-Exposure Fusion
Yao Liu, Lishen Qu, ... (+7 more)
|
|
cs.CV
|
0 |
28 days ago |
| 219 |
SafeGuard: A Multi-Agent Perception-Reasoning Framework for Social-Risk AI-Generated Video Detection
Wenlin Wu, Sheng Zhou, ... (+4 more)
|
|
cs.CV
|
0 |
28 days ago |
| 220 |
Natural Language Camera Movement Understanding
Yuwen Tan, Joey Huang, ... (+3 more)
|
|
cs.CV
|
0 |
28 days ago |
| 221 |
$C^3$ASD: Multi-Level Consistency-Driven Representation Learning
Jin Hong, Jisoo Park, Junseok Kwon
|
|
cs.CV
|
0 |
28 days ago |
| 222 |
GuideMe: Multi-Domain Task Guidance and Intervention in Streaming Video
Fang Liu, Jinpeng Chen, ... (+8 more)
|
|
cs.CV
|
0 |
28 days ago |
| 223 |
Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning
Wenzheng Zeng, Siyi Jiao, ... (+3 more)
|
|
cs.CV
|
0 |
28 days ago |
| 224 |
Incentivizing Vision Language Models to Search for Long Video Question Answering
Harsh Goel, S P Sharan, ... (+5 more)
|
|
cs.CV
|
0 |
28 days ago |
| 225 |
Holo-Captioning: Toward the Text Equivalent of 3D Scenes
Kun-Yu Lin, Chengke Bu, ... (+2 more)
|
|
cs.CV
|
0 |
28 days ago |
| 226 |
ProLaViT: Learning Progressive Latent Visual Thoughts in Structured Latent Space
Peiming Li, Yifan Wang, ... (+5 more)
|
|
cs.CV
|
0 |
28 days ago |
| 227 |
SPLIT: Training-Free AI-Generated and Partially Edited Video Detection via Spatial Patch-Level Incoherence and Temporal Roughness
Jongyeop Hyun, Hyounghun Kim
|
|
cs.CV
|
0 |
28 days ago |
| 228 |
Less Tokens, Better Forecasts: Sparse Residual Routing for Efficient Weather Prediction
Janet Wang, Yunbei Zhang, ... (+4 more)
|
|
cs.LG
|
0 |
29 days ago |
| 229 |
Conversational Human Audio-visual Talking Dialogue Generation
Junhao Song, Lluis Guasch, ... (+9 more)
|
|
cs.CV
|
0 |
29 days ago |
| 230 |
Global Pose Control for Generative View Synthesis in Normalized Object Coordinate Space
Zhibing Li, Amogh Gupta, ... (+2 more)
|
|
cs.CV
|
0 |
29 days ago |
| 231 |
Seek to Segment: Active Perception for Panoramic Referring Segmentation
Song Tang, Shuming Hu, ... (+3 more)
|
|
cs.CV
|
0 |
29 days ago |
| 232 |
Towards Robustness against Typographic Attack with Training-free Concept Localization
Bohan Liu, Wenqian Ye, ... (+4 more)
|
|
cs.CV
|
0 |
29 days ago |
| 233 |
GeoMix: Descriptor-Free Visual Localization via Global Context and Multi-Detector Training
Yejun Zhang, Xinjue Wang, ... (+3 more)
|
|
cs.CV
|
0 |
29 days ago |
| 234 |
Combating Textual Noise and Redundancy: Entropy-Aware Dense Visual Token Pruning
Xuehui Wang, Xuankun Yang, Wei Shen
|
|
cs.CV
|
0 |
29 days ago |
| 235 |
BiSLW: Bi-Spectral Latent Watermarking for Generative Diffusion Models
Aryan Pandit
|
|
cs.CV
|
0 |
29 days ago |
| 236 |
Interpretation-Oriented Cloud Removal via Observation-Anchored Residual Flow with Geo-Contextual Alignment
Ziyao Wang, Maonan Wang, ... (+6 more)
|
|
cs.CV
|
0 |
29 days ago |
| 237 |
Wavelet-Guided Semantic Signal Compensation for Inversion-Free Image Editing
Anqi Tang, Wenhao Sun, Zhaoqiang Liu
|
|
cs.CV
|
0 |
29 days ago |
| 238 |
Learning Spectral and Polarimetric Clues for One-to-Multimodal Novel View Synthesis
Federico Lincetto, Gianluca Agresti, ... (+3 more)
|
|
cs.CV
|
0 |
29 days ago |
| 239 |
When Token Compression Breaks: Structural Pruning vs. Token Reduction for Robust ViT Segmentation under High Compression
Tien-Phat Nguyen, Ngai-Man Cheung
|
|
cs.CV
|
0 |
29 days ago |
| 240 |
LongEgoRefer: A Benchmark for Long-Form Egocentric Video Referring Expression Comprehension
Shunya Kato, Taiki Miyanishi, ... (+4 more)
|
|
cs.CV
|
0 |
29 days ago |
| 241 |
Comprehensive Robustness Analysis of LiDAR-based 3D Object Detection in Autonomous Driving
Adwait Chandorkar, Kai Krink, ... (+3 more)
|
|
cs.CV
|
0 |
29 days ago |
| 242 |
UnderOneFacade: Worldwide Facade Semantic Segmentation Benchmark Dataset
Yi Wang, Fan Wang, ... (+10 more)
|
|
cs.CV
|
0 |
29 days ago |
| 243 |
Training-free Controllable Human Motion Generation under Heterogeneous Constraints
Xiaofei Hui, Bo Yan, ... (+3 more)
|
|
cs.CV
|
0 |
29 days ago |
| 244 |
Understanding Geometric Representations in Self-Supervised Vision Transformers via Subspace Intervention
Weichen Zhou, Yawen Zou, ... (+4 more)
|
|
cs.CV
|
0 |
29 days ago |
| 245 |
Open-Weather Robust 3D Detection via Dual-Critic Diffusion Alignment
Shuyao Li, Chuanxing Geng, ... (+3 more)
|
|
cs.CV
|
0 |
29 days ago |
| 246 |
NeoMap: Training-free Novel-View Synthesis from Single Images and Videos
Jinxi Li, Tianyi Zhang, ... (+5 more)
|
|
cs.CV
|
0 |
29 days ago |
| 247 |
PhysMani: Physics-principled 3D World Model for Dynamic Object Manipulation
Peng Yun, Shouwang Huang, ... (+4 more)
|
|
cs.RO
|
0 |
29 days ago |
| 248 |
SFKD: Spatial--Frequency Joint-Aware Heterogeneous Knowledge Distillation via Multi-Level Wavelet Spectral Interaction
Cuipeng Wang, Haipeng Wang
|
|
cs.CV
|
0 |
29 days ago |
| 249 |
Diversity-aware View Partitioning for Scalable VGGT
Jinsoo Park, Donggyu Choi, ... (+3 more)
|
|
cs.CV
|
0 |
29 days ago |
| 250 |
Geometric Foundation Model Distillation for Efficient Lunar 3D Reconstruction
Clémentine Grethen, Florient Chouteau, ... (+2 more)
|
|
cs.CV
|
0 |
29 days ago |