6 papers
TSHA: A Benchmark for Visual Language Models in Trustworthy Safety Hazard Assessment Scenarios
Qiucheng Yu, Ruijie Xu, Mingang Chen +2
The paper introduces TSHA, a large benchmark of real-world indoor safety hazard assessment questions for evaluating vision‑language models, and shows that training on this data imp…
Rethinking Weakly-supervised Video Temporal Grounding From a Game Perspective
Xiang Fang, Zeyu Xiong, Wanlong Fang +7
This paper addresses the challenging task of weakly-supervised video temporal grounding. Existing approaches are generally based on the moment proposal selection framework that uti…
Audio Does Matter: Importance-Aware Multi-Granularity Fusion for Video Moment Retrieval
Junan Lin, Daizong Liu, Xianke Chen +5
Video Moment Retrieval (VMR) aims to retrieve a specific moment semantically related to the given query. To tackle this task, most existing VMR methods solely focus on the visual a…
Multimodal Language Models See Better When They Look Shallower
Haoran Chen, Junyan Lin, Xinghao Chen +6
Multimodal large language models (MLLMs) typically extract visual features from the final layers of a pretrained Vision Transformer (ViT). This widespread deep-layer bias, however,…
Towards Efficient General Feature Prediction in Masked Skeleton Modeling
Shengkai Sun, Zefan Zhang, Jianfeng Dong +3
Recent advances in the masked autoencoder (MAE) paradigm have significantly propelled self-supervised skeleton-based action recognition. However, most existing approaches limit rec…
UW-3DGS: Underwater 3D Reconstruction with Physics-Aware Gaussian Splatting
Wenpeng Xing, Jie Chen, Zaifeng Yang +5
Underwater 3D scene reconstruction faces severe challenges from light absorption, scattering, and turbidity, which degrade geometry and color fidelity in traditional methods like N…