| 251 |
C2E: Boosting Ego-Only 3D Object Detection via Multi-Teacher Contrastive Knowledge Distillation
Jinlong Wang, Xun Huang, ... (+3 more)
|
|
cs.CV
|
0 |
29 days ago |
| 252 |
PixGS: Pixel-Space Diffusion for Direct 3D Gaussian Splat Generation
Cao Duy, Phong Nguyen-Ha
|
|
cs.CV
|
0 |
29 days ago |
| 253 |
Path-level Hindsight Instructions for Semantic Exploration in Vision-Language Navigation
Sung June Kim, Sangpil Kim, Honglak Lee
|
|
cs.AI
|
0 |
29 days ago |
| 254 |
InterCMDM: Block-Causal Diffusion for Autoregressive Human Interaction Generation
Qing Yu, Kent Fujiwara
|
|
cs.CV
|
0 |
29 days ago |
| 255 |
ReQuest: Rethinking-based Question-Aware Frame Selection for Long-Form Video QA
Minkuk Kim, Suyong Yun, ... (+4 more)
|
|
cs.CV
|
0 |
29 days ago |
| 256 |
LASER: A Corrective Lens for LVLMs via Visual Attention Preservation and Sink Suppression
Bowen Yuan, Zijian Wang, ... (+3 more)
|
|
cs.CV
|
0 |
29 days ago |
| 257 |
ICDepth: Taming Video Diffusion Models for Video Depth Estimation via In-Context Conditioning
Xuanhua He, Jiaxin Xie, ... (+2 more)
|
|
cs.CV
|
0 |
29 days ago |
| 258 |
Temporal and Cross-Modal Alignment for Enhanced Audiovisual Video Captioning
Chen Zhao, Jiajun Ma, ... (+7 more)
|
|
cs.CV
|
0 |
29 days ago |
| 259 |
Unified Panoramic-Gaussian Representation for Monocular 4D Scene Synthesis
Yuankun Yang, Yi Wei, ... (+2 more)
|
|
cs.CV
|
0 |
29 days ago |
| 260 |
Teaching Vision-Language-Action Models What to See and Where to Look
Yuguang Yang, Canyu Chen, ... (+11 more)
|
|
cs.CV
|
0 |
29 days ago |
| 261 |
Domain Generalization via Text-Anchored Information Bottleneck
Eunyi Lyou, Yunjeong Choi, ... (+2 more)
|
|
cs.CV
|
0 |
29 days ago |
| 262 |
Disentangling Pictorial Cue Understanding from Language Bias in VLMs via Depth Ordering Task
Yiqian Liu, Iuliia Kotseruba, John K. Tsotsos
|
|
cs.CV
|
0 |
1 month ago |
| 263 |
Anti-Prompt: Image Protection against Text-Guided Image-to-Video Generation
Yeonghwan Song, Chanhui Lee, ... (+2 more)
|
|
cs.CV
|
0 |
1 month ago |
| 264 |
MapDreamer: Aerial Imagery Conditioned Latent Diffusion for Lane-Level Map Generation
Julian Brandes, Philipp Crocoll, Wolfram Burgard
|
|
cs.CV
|
0 |
1 month ago |
| 265 |
Ink3D: Sculpting 3D Assets with Extremely Complex Textures via Video Generative Models
Yue Han, Chong Li, ... (+6 more)
|
|
cs.CV
|
0 |
1 month ago |
| 266 |
Towards Metric-Agnostic Trajectory Forecasting
Markus Knoche, Daan de Geus, Bastian Leibe
|
|
cs.CV
|
0 |
1 month ago |
| 267 |
LongVQUBench: Benchmarking Long-Term Video Quality Understanding of Vision-Language Models
Arpita Nema, Hanwei Zhu, ... (+2 more)
|
|
cs.CV
|
0 |
1 month ago |
| 268 |
AutoSpeed: Annotation-Free Stage-Adaptive Motion Speed Learning for Robot Manipulation
Qingda Hu, Ziheng Qiu, ... (+3 more)
|
|
cs.RO
|
0 |
1 month ago |
| 269 |
AVSR-Diff: Scale-Agnostic Diffusion Priors for Temporally Consistent Arbitrary-Scale Video Super-Resolution
Geunhyuk Youk, Jeonghyeok Do, ... (+3 more)
|
|
cs.CV
|
0 |
1 month ago |
| 270 |
Condensing Large-Scale Datasets Directly with Minimal Information Loss
Xinyi Shang, Peng Sun, ... (+3 more)
|
|
cs.CV
|
0 |
1 month ago |
| 271 |
MG-RWKV: Multi-Grained Context-Aware RWKV for Temporal Forgery Localization
Jingchen Ni, Cangjin Yu, ... (+7 more)
|
|
cs.CV
|
0 |
1 month ago |
| 272 |
DeWorldSG: Depth-Aware 3D Semantic Scene Graph Generation via World-Model Priors
Seok-Young Kim, Abdelrahman Elskhawy, ... (+5 more)
|
|
cs.CV
|
0 |
1 month ago |
| 273 |
Improving Sparse-View 3DGS Generalization via Flat Minima Optimization
Kangmin Seo, Sangeek Hyun, ... (+2 more)
|
|
cs.CV
|
0 |
1 month ago |
| 274 |
MoVA: Learning Asymmetric Dual Projections for Modular Long Video-Text Alignment
Peiyuan Zhu, Shaoan Xie, ... (+6 more)
|
|
cs.CV
|
0 |
1 month ago |
| 275 |
GKDT: General Keypoint Detection Transformer
Changsheng Lu, Yuxin Chen, ... (+6 more)
|
|
cs.CV
|
0 |
1 month ago |
| 276 |
Towards Memory-Efficient Autoregressive Video Generation via Instance-Specific Parametric Absorption
Xiaomeng Fu, Jia Li, ... (+6 more)
|
|
cs.CV
|
0 |
1 month ago |
| 277 |
AdaBoosting Text Prompts for Vision-Language Models
Seokhee Jin, Changhwan Sung, ... (+3 more)
|
|
cs.LG
|
0 |
1 month ago |
| 278 |
Domain Arithmetic: One-Shot VLA Adaptation under Environmental Shifts
Taewook Kang, Taeheon Kim, ... (+2 more)
|
|
cs.RO
|
0 |
1 month ago |
| 279 |
Not All Prediction Targets Keep Training-Free Diffusion Guidance on the Manifold
Yunsung Lee, Hyeongmin Lee
|
|
cs.CV
|
0 |
1 month ago |
| 280 |
Token-level Response-visual Attention Guidance for Multimodal LLMs Knowledge Distillation
Jaehyun Jang, Eunseop Yoon, ... (+4 more)
|
|
cs.CV
|
0 |
1 month ago |
| 281 |
EPO: Boosting 3D Foundation Models with Edge-based Pose Optimization
Mattia D'Urso, Christian Sormann, ... (+2 more)
|
|
cs.CV
|
0 |
1 month ago |
| 282 |
Caption Bottleneck Models
Seref Baris Cagliyan, Umut Ozdemir, ... (+2 more)
|
|
cs.CV
|
0 |
1 month ago |
| 283 |
BrainFIBRE: A Foundation Model via Information Decomposition for Brain Microstructure
Zijian Dong, Yi Lin, ... (+4 more)
|
|
cs.CV
|
0 |
1 month ago |
| 284 |
SPECSIA: Stylization Dataset for Novel-View Enhancement in Drawing-based 3D Animation
Kyuwon Kim, Sunjae Yoon, Chang D. Yoo
|
|
cs.CV
|
0 |
1 month ago |
| 285 |
HieDG: A Hierarchical Discrete Geometry-Guided Framework for Multi-Animal Tracking
Chenxun Deng, Zhongde Zhang, ... (+8 more)
|
|
cs.CV
|
0 |
1 month ago |
| 286 |
GenSP: Consistent Spherical Parameterization via Learning Shape Generative Models
Sai Karthikey Pentapati, Shashank Gupta, ... (+4 more)
|
|
cs.CV
|
0 |
1 month ago |
| 287 |
StochasT: Learning with Stochastic Turn Depth for Visual Instruction Tuning
Yuan Qing, Chengzhi Mao, Boqing Gong
|
|
cs.CV
|
0 |
1 month ago |
| 288 |
Multi-scale Mixture of World Models for Embodied Agents in Evolving Environments
Jinwoo Jang, Daniel J. Rho, ... (+3 more)
|
|
cs.AI
|
0 |
1 month ago |
| 289 |
VideoSearch-R1: Iterative Video Retrieval and Reasoning via Soft Query Refinement
Seohyun Lee, Seoung Choi, ... (+3 more)
|
|
cs.CV
|
0 |
1 month ago |
| 290 |
Information-Regularized Attention for Visual-Centric Reasoning
Guohao Sun, Xiaofang Wang, ... (+4 more)
|
|
cs.CV
|
0 |
1 month ago |
| 291 |
HyFL-CLIP: Hyperbolic Fine-Tuning of CLIP for Robust Long-Context Understanding
Ji Ha Jang, Hayeon Kim, ... (+3 more)
|
|
cs.CV
|
0 |
1 month ago |
| 292 |
MedCAGD: Context-Aware Gated Decoder for Efficient Medical Image Segmentation
Saad Wazir, Patrick Dominique Vibild, ... (+3 more)
|
|
cs.CV
|
0 |
1 month ago |
| 293 |
Personalization as Inverse Planning: Learning Latent Design Intents for Agentic Slide Generation via Structural Denoising
Tianci Liu, Zihan Dong, ... (+6 more)
|
|
cs.AI
|
0 |
1 month ago |
| 294 |
The Illusion of High Utility in Safety Alignment of Text-to-Image Diffusion Models
Adeel Yousaf, Soumik Ghosh, ... (+3 more)
|
|
cs.CV
|
0 |
1 month ago |
| 295 |
Vitality-Aware Compression for Efficient Image-to-Shape Diffusion Transformers
Jaeah Lee, Hyunjin Kim, ... (+2 more)
|
|
cs.CV
|
0 |
1 month ago |
| 296 |
Learning to Compose: Revisiting Proxy Task Design for Zero-Shot Composed Image Retrieval
Jingjing Zhang, Lei Zhang, ... (+2 more)
|
|
cs.CV
|
0 |
1 month ago |
| 297 |
MEPA: Multi-Scale Representation Alignment for Visual Autoregressive Modeling with Mixture of Experts
Nuoyan Zhou, Zhijun Tu, ... (+5 more)
|
|
cs.CV
|
0 |
1 month ago |
| 298 |
SFDATrack: Generalized Source-Free Domain Adaptive Tracking Under Adverse Weather Conditions
Siyuan Yao, Ziqi Wang, ... (+4 more)
|
|
cs.CV
|
0 |
1 month ago |
| 299 |
ReShift: Aha-Moment-Driven Reasoning-Level Backdoor Attacks on Vision-Language Models
Zhihao Dou, Qinjian Zhao, ... (+2 more)
|
|
cs.CR
|
0 |
1 month ago |
| 300 |
Wake up for Touch! Mask-isolated Tactile Alignment Learning in MLLMs
Yoonhyung Park, Minji Kim, ... (+2 more)
|
|
cs.CV
|
0 |
1 month ago |