5 papers
FALCON: Future-Aware Learning with Contextual Object-Centric Pretraining for UAV Action Recognition
Ruiqi Xian, Xiyang Wu, Tianrui Guan +3
We introduce FALCON, a unified self-supervised video pretraining approach for UAV action recognition from raw RGB aerial footage, requiring no additional preprocessing at inference…
EgoSocial: Benchmarking Proactive Intervention Ability of Omnimodal LLMs via Egocentric Social Interaction Perception
Xijun Wang, Tanay Sharma, Achin Kulshrestha +3
As AR/VR technologies become integral to daily life, there's a growing need for AI that understands human social dynamics from an egocentric perspective. However, current LLMs ofte…
Bi-VLM: Pushing Ultra-Low Precision Post-Training Quantization Boundaries in Vision-Language Models
Xijun Wang, Junyun Huang, Rayyan Abdalla +3
We address the critical gap between the computational demands of vision-language models and the possible ultra-low-bit weight precision (bitwidth bits) we can use for highe…
AutoHallusion: Automatic Generation of Hallucination Benchmarks for Vision-Language Models
Xiyang Wu, Tianrui Guan, Dianqi Li +9
Large vision-language models (LVLMs) are prone to hallucinations, where certain contextual cues in an image can trigger the language module to produce overconfident and incorrect r…
AGL-NET: Aerial-Ground Cross-Modal Global Localization with Varying Scales
Tianrui Guan, Ruiqi Xian, Xijun Wang +4
We present AGL-NET, a novel learning-based method for global localization using LiDAR point clouds and satellite maps. AGL-NET tackles two critical challenges: bridging the represe…