works on

From the 1 of 9 linked papers with an AI index.

activity
20242026
collaborators

9 papers

cs.CV2026

Parse, Search, and Confirmation: Training-Free Aerial Vision-and-Dialog Navigation with Chain-of-Thought Reasoning and Structured Spatial Memory

Yu Qi, Hongyu Li, Shaofei Huang +6

The paper introduces PSC-AVDN, a training‑free framework for high‑altitude UAV navigation that parses dialog instructions, searches with chain‑of‑thought reasoning, and confirms ta…

cs.CV2026

Temporal and Cross-Modal Alignment for Enhanced Audiovisual Video Captioning

Chen Zhao, Jiajun Ma, Qilong Huang +6

While Multimodal Large Language Models (MLLMs) have advanced video understanding, achieving precise temporal and cross-modal alignment in audiovisual video captioning remains a for…

cs.CV2026

Video-MME-Logical: A Controlled Diagnostic Benchmark for Video Temporal-Logical Reasoning

Hohin Kwan, Hongyu Li, Ray Zhang +5

Recent interest in multimodal large language models (MLLMs) raises a central question: can they reason over dynamic visual evidence rather than merely recognize objects or events i…

cs.RO2026

From Instruction to Event: Sound-Triggered Mobile Manipulation

Hao Ju, Shaofei Huang, Hongyu Li +4

Current mobile manipulation research predominantly follows an instruction-driven paradigm, where agents rely on predefined textual commands to execute tasks. However, this setting…

cs.CV2026

RPiAE: A Representation-Pivoted Autoencoder Enhancing Both Image Generation and Editing

Yue Gong, Hongyu Li, Shanyuan Liu +8

Diffusion models have become the dominant paradigm for image generation and editing, with latent diffusion models shifting denoising to a compact latent space for efficiency and sc…

cs.CV2025

Revisiting Audio-Visual Segmentation with Vision-Centric Transformer

Shaofei Huang, Rui Ling, Tianrui Hui +6

Audio-Visual Segmentation (AVS) aims to segment sound-producing objects in video frames based on the associated audio signal. Prevailing AVS methods typically adopt an audio-centri…