activity
20242026
collaborators

6 papers

cs.CV2026

FALCON: Future-Aware Learning with Contextual Object-Centric Pretraining for UAV Action Recognition

Ruiqi Xian, Xiyang Wu, Tianrui Guan +3

We introduce FALCON, a unified self-supervised video pretraining approach for UAV action recognition from raw RGB aerial footage, requiring no additional preprocessing at inference…

cs.CV2025

CaughtCheating: Is Your MLLM a Good Cheating Detective? Exploring the Boundary of Visual Perception and Reasoning

Ming Li, Chenguang Wang, Yijun Liang +6

Recent agentic Multi-Modal Large Language Models (MLLMs) such as GPT-o3 have achieved near-ceiling scores on various existing benchmarks, motivating a demand for more challenging t…

cs.RO2025

On the Vulnerability of LLM/VLM-Controlled Robotics

Xiyang Wu, Souradip Chakraborty, Ruiqi Xian +6

In this work, we highlight vulnerabilities in robotic systems integrating large language models (LLMs) and vision-language models (VLMs) due to input modality sensitivities. While…

cs.CV2024

AutoHallusion: Automatic Generation of Hallucination Benchmarks for Vision-Language Models

Xiyang Wu, Tianrui Guan, Dianqi Li +9

Large vision-language models (LVLMs) are prone to hallucinations, where certain contextual cues in an image can trigger the language module to produce overconfident and incorrect r…

cs.CV2024

AGL-NET: Aerial-Ground Cross-Modal Global Localization with Varying Scales

Tianrui Guan, Ruiqi Xian, Xijun Wang +4

We present AGL-NET, a novel learning-based method for global localization using LiDAR point clouds and satellite maps. AGL-NET tackles two critical challenges: bridging the represe…

cs.RO2024

LANCAR: Leveraging Language for Context-Aware Robot Locomotion in Unstructured Environments

Chak Lam Shek, Xiyang Wu, Wesley A. Suttle +5

Navigating robots through unstructured terrains is challenging, primarily due to the dynamic environmental changes. While humans adeptly navigate such terrains by using context fro…