| 151 |
Latent Space Reinforcement Learning for Inverse Material Estimation in Food Fracture Simulation
Adrian Ramlal, Yuhao Chen, John S. Zelek
|
|
cs.CV
|
0 |
1 month ago |
| 152 |
VinQA: Visual Elements Interleaved Long-form Answer Generation for Real-World Multimodal Document QA
Young Rok Jang, Hyesoo Kong, ... (+4 more)
|
|
cs.CV
|
0 |
1 month ago |
| 153 |
Post-Launch Capability Expansion of Vision-Language Models via Prompting for On-Orbit Spacecraft Inspection
Nicholas A. Welsh, Lennon J. Shikhman, ... (+4 more)
|
|
cs.LG
|
0 |
1 month ago |
| 154 |
One Layer's Trash is Another Layer's Treasure: Adaptive Layer-wise Visual Token Selection in LVLMs
Yongru Chen, Kai Zhang, ... (+5 more)
|
|
cs.CV
|
0 |
1 month ago |
| 155 |
FloVerse: Floor Plan-Guided Multi-Modal Navigation
Weiqi Huang, Shuangyi Dong, ... (+4 more)
|
|
cs.RO
|
0 |
1 month ago |
| 156 |
Reasoning for Mobile User Experience with Multimodal LLMs: Task, Benchmark, and Approach
Ruichao Mao, Zhou Fang, ... (+10 more)
|
|
cs.AI
|
0 |
1 month ago |
| 157 |
Context-Aware Feature-Fusion for Co-occurring Object Detection in Autonomous Driving
Binay Kumar Singh, Niels Da Vitoria Lobo
|
|
cs.CV
|
0 |
1 month ago |
| 158 |
Metadata-Aware Multi-Prompt Reasoning for Zero-Shot Accident Understanding
Tarandeep Singh, Soumyanetra Pal, ... (+2 more)
|
|
cs.CV
|
0 |
1 month ago |
| 159 |
AutoMine Solution for AV2 2026 Scenario Mining Challenge
Songliang Cao, Jiele Zhao, ... (+11 more)
|
|
cs.AI
|
0 |
1 month ago |
| 160 |
Information-Theoretic Decomposition for Multimodal Interaction Learning
Zequn Yang, Yake Wei, ... (+3 more)
|
|
cs.LG
|
0 |
1 month ago |
| 161 |
MOFA-VTON: More Fashion Possibilities with Fine-Grained Adaptations in Virtual Try-On
Xiaoyu Han, Chenyang Wang, ... (+4 more)
|
|
cs.CV
|
0 |
1 month ago |
| 162 |
Monte Carlo Pass Search: Using Trajectory Generation for 3D Counterfactual Pass Evaluation in Football
Andrew Kang, Priya Narasimhan
|
|
cs.AI
|
0 |
1 month ago |
| 163 |
DB-3DME: From Dataset to Benchmark for Human-aligned Automatic 3D Mesh Evaluation
Nanshan Jia, Zhenyu Zhao, ... (+3 more)
|
|
cs.CV
|
0 |
1 month ago |
| 164 |
The Impact of VAE Design on Latent Pose Representations for Diffusion-based Sign Language Production
Guilhem Fauré, Mostafa Sadeghi, ... (+2 more)
|
|
cs.AI
|
0 |
29 days ago |
| 165 |
Reference-Free Assessment of Physical Consistency in World Model-based Video Generation
Yun Oh, Sukmin Yun
|
|
cs.AI
|
0 |
1 month ago |
| 166 |
Zero-Shot Vision-Language Models for Classroom Engagement Recognition: A Benchmark Study of Prompt Sensitivity and Cross-Dataset Generalization
Aman Goyal, Kshama Nitin Shah, Kemmannu Vineet Venkatesh Rao
|
|
cs.CV
|
0 |
1 month ago |
| 167 |
Beyond Flat Labels: Level-Restricted Contrastive Learning for Hierarchical Fine-Grained Vision Classification
Zhiyuan Tao, Srikumar Sastry, ... (+10 more)
|
|
cs.CV
|
0 |
1 month ago |
| 168 |
Scene-Level Heterogeneous Physics Simulation with 3D Gaussian Splats
Xiaoyang Liu, Shangzhe Wu, Kai Han
|
|
cs.GR
|
0 |
1 month ago |
| 169 |
Adversarial Domain Prompt Tuning and Generation for Single Domain Generalization
Zhipeng Xu, De Cheng, ... (+4 more)
|
|
cs.CV
|
0 |
1 month ago |
| 170 |
BIT-Nav: Brain-Inspired Trajectory Memory for Embodied Navigation
Rithvik Jonna, Aakash Gurram, ... (+3 more)
|
|
cs.RO
|
0 |
1 month ago |
| 171 |
Linear Recurrent Unit with Semantic Modulation for Image Super-Resolution
Mingyu Choi, Woo Kyoung Han, ... (+2 more)
|
|
cs.CV
|
0 |
1 month ago |
| 172 |
SCR-Guided Difficulty-Aware Optimization for Infrared Small Target Detection
Yunus Sevim, Behçet Uğur Töreyin
|
|
cs.CV
|
0 |
1 month ago |
| 173 |
InTrain: Intrinsic Trainability for Zero-Cost Neural Architecture Search
Qinqin Zhou, Fuhai Chen, ... (+4 more)
|
|
cs.LG
|
0 |
1 month ago |
| 174 |
SIR: Structured Image Representations for Explainable Robot Learning
Paul Mattes, Jan Schwab, ... (+6 more)
|
|
cs.RO
|
0 |
22 days ago |
| 175 |
Enhancing Part-Level Point Grounding for Any Open-Source MLLMs
Jin-Cheng Jhang, Fu-En Wang, ... (+5 more)
|
|
cs.CV
|
0 |
23 days ago |
| 176 |
TrafficAlign: Aligning Large Language Models for Traffic Scenario Generation
Zhi Tu, Liangkun Niu, Tianyi Zhang
|
|
cs.CV
|
0 |
24 days ago |
| 177 |
IMU-HOI: A Symbiotic Framework for Coherent Human-Object Interaction and Motion Capture via Contact-Conscious Inertial Fusion
Lizhou Lin, Songpengcheng Xia, ... (+4 more)
|
|
cs.CV
|
0 |
25 days ago |
| 178 |
Perceptual 3D Simulation With Physical World Modeling
Wanhee Lee, Klemen Kotar, ... (+3 more)
|
|
cs.CV
|
0 |
26 days ago |
| 179 |
Dataset Usage Inference without Shadow Models or Held-out Data
Wojciech Łapacz, Stanisław Pawlak, ... (+3 more)
|
|
cs.LG
|
0 |
27 days ago |
| 180 |
GRAFT: Graph-Based Affordance Transfer via Part Correspondence
Mengying Lin, Utkarsh Mishra, ... (+2 more)
|
|
cs.RO
|
0 |
28 days ago |
| 181 |
Universal Guideline-Driven Image Clustering via a Hybrid LLM Agent
Wenliang Zhong, Rob Barton, ... (+8 more)
|
|
cs.CV
|
0 |
28 days ago |
| 182 |
Mirror Illusion Art
Xiaopei Zhu, Zeyuan Li, ... (+2 more)
|
|
cs.CV
|
0 |
19 days ago |
| 183 |
How Much Future Helps? A Controlled Study of Future-Privileged Supervision for Causal Egocentric Gaze Estimation
Jia Li, Wenjie Zhao, ... (+6 more)
|
|
cs.CV
|
0 |
20 days ago |
| 184 |
Local Motion Matters: A Deconstruct-Recompose Paradigm for Reinforcement Learning Pre-training from Videos
Jinwen Wang, Youfang Lin, ... (+3 more)
|
|
cs.LG
|
0 |
20 days ago |
| 185 |
Phantom: A Unified Face-Swap Deepfake Protection Framework with Latent and Spatial Constraints
Jungkon Kim, Cheolseung Jung, ... (+2 more)
|
|
cs.CV
|
0 |
21 days ago |
| 186 |
Parse, Search, and Confirmation: Training-Free Aerial Vision-and-Dialog Navigation with Chain-of-Thought Reasoning and Structured Spatial Memory
Yu Qi, Hongyu Li, ... (+7 more)
|
|
cs.CV
|
0 |
8 days ago |
| 187 |
GDP.pdf: Benchmarking Grounded Multimodal Reasoning over Professional PDF Documents
Suhaas Garre, Emily Ritchie, ... (+2 more)
|
|
cs.CV
|
0 |
8 days ago |
| 188 |
SynCLIP: Synonym-Coherent Language-Image Pretraining for Robust Open-Vocabulary Dense Perception
Mingjie Xie, Guangjun He, ... (+6 more)
|
|
cs.CV
|
0 |
8 days ago |
| 189 |
GIRAF: Towards Generalizable Human Interactions with Articulated Objects
Xiaohan Zhang, Sebastian Starke, ... (+4 more)
|
|
cs.CV
|
0 |
13 days ago |
| 190 |
Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF
Eric Zhu, Abhinav Shrivastava, Soumik Mukhopadhyay
|
|
cs.LG
|
0 |
13 days ago |
| 191 |
Why Fake ? Unveiling the Semantic Vocabulary of Deepfake Detectors
Vazgken Vanian, Alexandros Doumanoglou, Dimitris Zarpalas
|
|
cs.CV
|
0 |
13 days ago |
| 192 |
Andha-Dhun: A First Look at Audio Descriptions in Hindi
Ritabrata Chakraborty, Divy Kala, ... (+4 more)
|
|
cs.CV
|
0 |
14 days ago |
| 193 |
Breaking Spurious Correlations via Generative Randomization and Cross-Variant Self-Supervised Learning
Suraj Yadav, Anjaneya Sharma, Siddharth Yadav
|
|
cs.CV
|
0 |
14 days ago |