From the 1 of 9 linked papers with an AI index.
9 papers
Parse, Search, and Confirmation: Training-Free Aerial Vision-and-Dialog Navigation with Chain-of-Thought Reasoning and Structured Spatial Memory
Yu Qi, Hongyu Li, Shaofei Huang +6
The paper introduces PSC-AVDN, a training‑free framework for high‑altitude UAV navigation that parses dialog instructions, searches with chain‑of‑thought reasoning, and confirms ta…
Temporal and Cross-Modal Alignment for Enhanced Audiovisual Video Captioning
Chen Zhao, Jiajun Ma, Qilong Huang +6
While Multimodal Large Language Models (MLLMs) have advanced video understanding, achieving precise temporal and cross-modal alignment in audiovisual video captioning remains a for…
Video-MME-Logical: A Controlled Diagnostic Benchmark for Video Temporal-Logical Reasoning
Hohin Kwan, Hongyu Li, Ray Zhang +5
Recent interest in multimodal large language models (MLLMs) raises a central question: can they reason over dynamic visual evidence rather than merely recognize objects or events i…
From Instruction to Event: Sound-Triggered Mobile Manipulation
Hao Ju, Shaofei Huang, Hongyu Li +4
Current mobile manipulation research predominantly follows an instruction-driven paradigm, where agents rely on predefined textual commands to execute tasks. However, this setting…
RPiAE: A Representation-Pivoted Autoencoder Enhancing Both Image Generation and Editing
Yue Gong, Hongyu Li, Shanyuan Liu +8
Diffusion models have become the dominant paradigm for image generation and editing, with latent diffusion models shifting denoising to a compact latent space for efficiency and sc…
Revisiting Audio-Visual Segmentation with Vision-Centric Transformer
Shaofei Huang, Rui Ling, Tianrui Hui +6
Audio-Visual Segmentation (AVS) aims to segment sound-producing objects in video frames based on the associated audio signal. Prevailing AVS methods typically adopt an audio-centri…