| 1 |
SoccerNet 2026 Player-Centric Ball-Action Spotting:Retraining and Post-Processing Extensions to the FOOTPASS Baselines
Parthsarthi Rawat
|
|
cs.CV
|
0 |
2 months ago |
| 2 |
EditSSC: Toward Editable Semantic Occupancy Scenes with Unconditional Diffusion Models
Fatima Balde, Raoul de Charette, Alexandre Boulch
|
|
cs.CV
|
0 |
2 months ago |
| 3 |
Latent Space Reinforcement Learning for Inverse Material Estimation in Food Fracture Simulation
Adrian Ramlal, Yuhao Chen, John S. Zelek
|
|
cs.CV
|
0 |
2 months ago |
| 4 |
VinQA: Visual Elements Interleaved Long-form Answer Generation for Real-World Multimodal Document QA
Young Rok Jang, Hyesoo Kong, ... (+4 more)
|
|
cs.CV
|
0 |
2 months ago |
| 5 |
Post-Launch Capability Expansion of Vision-Language Models via Prompting for On-Orbit Spacecraft Inspection
Nicholas A. Welsh, Lennon J. Shikhman, ... (+4 more)
|
|
cs.LG
|
0 |
2 months ago |
| 6 |
One Layer's Trash is Another Layer's Treasure: Adaptive Layer-wise Visual Token Selection in LVLMs
Yongru Chen, Kai Zhang, ... (+5 more)
|
|
cs.CV
|
0 |
2 months ago |
| 7 |
FloVerse: Floor Plan-Guided Multi-Modal Navigation
Weiqi Huang, Shuangyi Dong, ... (+4 more)
|
|
cs.RO
|
0 |
2 months ago |
| 8 |
Reasoning for Mobile User Experience with Multimodal LLMs: Task, Benchmark, and Approach
Ruichao Mao, Zhou Fang, ... (+10 more)
|
|
cs.AI
|
0 |
2 months ago |
| 9 |
Context-Aware Feature-Fusion for Co-occurring Object Detection in Autonomous Driving
Binay Kumar Singh, Niels Da Vitoria Lobo
|
|
cs.CV
|
0 |
2 months ago |
| 10 |
Metadata-Aware Multi-Prompt Reasoning for Zero-Shot Accident Understanding
Tarandeep Singh, Soumyanetra Pal, ... (+2 more)
|
|
cs.CV
|
0 |
2 months ago |
| 11 |
AutoMine Solution for AV2 2026 Scenario Mining Challenge
Songliang Cao, Jiele Zhao, ... (+11 more)
|
|
cs.AI
|
0 |
2 months ago |
| 12 |
Information-Theoretic Decomposition for Multimodal Interaction Learning
Zequn Yang, Yake Wei, ... (+3 more)
|
|
cs.LG
|
0 |
2 months ago |
| 13 |
MOFA-VTON: More Fashion Possibilities with Fine-Grained Adaptations in Virtual Try-On
Xiaoyu Han, Chenyang Wang, ... (+4 more)
|
|
cs.CV
|
0 |
2 months ago |
| 14 |
Monte Carlo Pass Search: Using Trajectory Generation for 3D Counterfactual Pass Evaluation in Football
Andrew Kang, Priya Narasimhan
|
|
cs.AI
|
0 |
2 months ago |
| 15 |
DB-3DME: From Dataset to Benchmark for Human-aligned Automatic 3D Mesh Evaluation
Nanshan Jia, Zhenyu Zhao, ... (+3 more)
|
|
cs.CV
|
0 |
2 months ago |
| 16 |
The Impact of VAE Design on Latent Pose Representations for Diffusion-based Sign Language Production
Guilhem Fauré, Mostafa Sadeghi, ... (+2 more)
|
|
cs.AI
|
0 |
2 months ago |
| 17 |
Reference-Free Assessment of Physical Consistency in World Model-based Video Generation
Yun Oh, Sukmin Yun
|
|
cs.AI
|
0 |
2 months ago |
| 18 |
Zero-Shot Vision-Language Models for Classroom Engagement Recognition: A Benchmark Study of Prompt Sensitivity and Cross-Dataset Generalization
Aman Goyal, Kshama Nitin Shah, Kemmannu Vineet Venkatesh Rao
|
|
cs.CV
|
0 |
2 months ago |
| 19 |
Beyond Flat Labels: Level-Restricted Contrastive Learning for Hierarchical Fine-Grained Vision Classification
Zhiyuan Tao, Srikumar Sastry, ... (+10 more)
|
|
cs.CV
|
0 |
2 months ago |
| 20 |
Scene-Level Heterogeneous Physics Simulation with 3D Gaussian Splats
Xiaoyang Liu, Shangzhe Wu, Kai Han
|
|
cs.GR
|
0 |
2 months ago |
| 21 |
Adversarial Domain Prompt Tuning and Generation for Single Domain Generalization
Zhipeng Xu, De Cheng, ... (+4 more)
|
|
cs.CV
|
0 |
2 months ago |
| 22 |
BIT-Nav: Brain-Inspired Trajectory Memory for Embodied Navigation
Rithvik Jonna, Aakash Gurram, ... (+3 more)
|
|
cs.RO
|
0 |
2 months ago |
| 23 |
Linear Recurrent Unit with Semantic Modulation for Image Super-Resolution
Mingyu Choi, Woo Kyoung Han, ... (+2 more)
|
|
cs.CV
|
0 |
2 months ago |
| 24 |
SCR-Guided Difficulty-Aware Optimization for Infrared Small Target Detection
Yunus Sevim, Behçet Uğur Töreyin
|
|
cs.CV
|
0 |
2 months ago |
| 25 |
InTrain: Intrinsic Trainability for Zero-Cost Neural Architecture Search
Qinqin Zhou, Fuhai Chen, ... (+4 more)
|
|
cs.LG
|
0 |
2 months ago |
| 26 |
SIR: Structured Image Representations for Explainable Robot Learning
Paul Mattes, Jan Schwab, ... (+6 more)
|
|
cs.RO
|
0 |
2 months ago |
| 27 |
Enhancing Part-Level Point Grounding for Any Open-Source MLLMs
Jin-Cheng Jhang, Fu-En Wang, ... (+5 more)
|
|
cs.CV
|
0 |
2 months ago |
| 28 |
TrafficAlign: Aligning Large Language Models for Traffic Scenario Generation
Zhi Tu, Liangkun Niu, Tianyi Zhang
|
|
cs.CV
|
0 |
2 months ago |
| 29 |
IMU-HOI: A Symbiotic Framework for Coherent Human-Object Interaction and Motion Capture via Contact-Conscious Inertial Fusion
Lizhou Lin, Songpengcheng Xia, ... (+4 more)
|
|
cs.CV
|
0 |
2 months ago |
| 30 |
Perceptual 3D Simulation With Physical World Modeling
Wanhee Lee, Klemen Kotar, ... (+3 more)
|
|
cs.CV
|
0 |
2 months ago |
| 31 |
Dataset Usage Inference without Shadow Models or Held-out Data
Wojciech Łapacz, Stanisław Pawlak, ... (+3 more)
|
|
cs.LG
|
0 |
2 months ago |
| 32 |
GRAFT: Graph-Based Affordance Transfer via Part Correspondence
Mengying Lin, Utkarsh Mishra, ... (+2 more)
|
|
cs.RO
|
0 |
2 months ago |
| 33 |
Universal Guideline-Driven Image Clustering via a Hybrid LLM Agent
Wenliang Zhong, Rob Barton, ... (+8 more)
|
|
cs.CV
|
0 |
2 months ago |
| 34 |
Mirror Illusion Art
Xiaopei Zhu, Zeyuan Li, ... (+2 more)
|
|
cs.CV
|
0 |
2 months ago |
| 35 |
How Much Future Helps? A Controlled Study of Future-Privileged Supervision for Causal Egocentric Gaze Estimation
Jia Li, Wenjie Zhao, ... (+6 more)
|
|
cs.CV
|
0 |
2 months ago |
| 36 |
Local Motion Matters: A Deconstruct-Recompose Paradigm for Reinforcement Learning Pre-training from Videos
Jinwen Wang, Youfang Lin, ... (+3 more)
|
|
cs.LG
|
0 |
2 months ago |
| 37 |
Phantom: A Unified Face-Swap Deepfake Protection Framework with Latent and Spatial Constraints
Jungkon Kim, Cheolseung Jung, ... (+2 more)
|
|
cs.CV
|
0 |
2 months ago |
| 38 |
Parse, Search, and Confirmation: Training-Free Aerial Vision-and-Dialog Navigation with Chain-of-Thought Reasoning and Structured Spatial Memory
Yu Qi, Hongyu Li, ... (+7 more)
|
|
cs.CV
|
0 |
1 month ago |
| 39 |
GDP.pdf: Benchmarking Grounded Multimodal Reasoning over Professional PDF Documents
Suhaas Garre, Emily Ritchie, ... (+2 more)
|
|
cs.CV
|
0 |
1 month ago |
| 40 |
SynCLIP: Synonym-Coherent Language-Image Pretraining for Robust Open-Vocabulary Dense Perception
Mingjie Xie, Guangjun He, ... (+6 more)
|
|
cs.CV
|
0 |
1 month ago |
| 41 |
GIRAF: Towards Generalizable Human Interactions with Articulated Objects
Xiaohan Zhang, Sebastian Starke, ... (+4 more)
|
|
cs.CV
|
0 |
1 month ago |
| 42 |
Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF
Eric Zhu, Abhinav Shrivastava, Soumik Mukhopadhyay
|
|
cs.LG
|
0 |
1 month ago |
| 43 |
Why Fake ? Unveiling the Semantic Vocabulary of Deepfake Detectors
Vazgken Vanian, Alexandros Doumanoglou, Dimitris Zarpalas
|
|
cs.CV
|
0 |
1 month ago |
| 44 |
Andha-Dhun: A First Look at Audio Descriptions in Hindi
Ritabrata Chakraborty, Divy Kala, ... (+4 more)
|
|
cs.CV
|
0 |
2 months ago |
| 45 |
Breaking Spurious Correlations via Generative Randomization and Cross-Variant Self-Supervised Learning
Suraj Yadav, Anjaneya Sharma, Siddharth Yadav
|
|
cs.CV
|
0 |
2 months ago |
| 46 |
CaT-GS: Efficient 3DGS Rendering for Large Scale Scenes via Inter-frame Caching and Tile Scheduling
Tingjia Zhang, Bo Chen, ... (+3 more)
|
|
cs.CV
|
0 |
1 month ago |
| 47 |
Test-Time Scaling for World Action Models via Zero-Shot Geometric Evaluation
Zesen Zhao, Minkyoung Cho, ... (+5 more)
|
|
cs.RO
|
0 |
1 month ago |
| 48 |
Illuminating Visual Identity in Universal Multimodal Embeddings
Jiawei Cao, Junyi Feng, ... (+6 more)
|
|
cs.CV
|
0 |
1 month ago |
| 49 |
The 1st AI Children Challenge
Boyi Li, Yifan Shen, ... (+8 more)
|
|
cs.CV
|
0 |
1 month ago |
| 50 |
Objective-Aligned Direct Answer SFT for Robust Multi-Frame Medical VQA
Site Li, Jianyi Hao, Xiaofeng Liu
|
|
cs.CV
|
0 |
1 month ago |