collaborators

6 papers

cs.MM2026

Hierarchical Semantic Correlation-Aware Masked Autoencoder for Unsupervised Audio-Visual Representation Learning

Donghuo Zeng, Hao Niu, Masato Taya

Learning aligned multimodal embeddings from weakly paired, label-free corpora is challenging: pipelines often provide only pre-extracted features, clips contain multiple events, an…

cs.CV2026

Variance & Greediness: A comparative study of metric-learning losses

Donghuo Zeng, Hao Niu, Zhi Li +1

Metric learning is central to retrieval, yet its effects on embedding geometry and optimization dynamics are not well understood. We introduce a diagnostic framework, VARIANCE (int…

cs.MM2026

Learning Audio-Visual Embeddings with Inferred Latent Interaction Graphs

Donghuo Zeng, Hao Niu, Yanan Wang +1

Learning robust audio-visual embeddings requires bringing genuinely related audio and visual signals together while filtering out incidental co-occurrences - background noise, unre…

cs.IR2025

An Empirical Study for Representations of Videos in Video Question Answering via MLLMs

Zhi Li, Yanan Wang, Hao Niu +2

Multimodal large language models have recently achieved remarkable progress in video question answering (VideoQA) by jointly processing visual, textual, and audio information. Howe…

cs.CV2025

CoTasks: Chain-of-Thought based Video Instruction Tuning Tasks

Yanan Wang, Julio Vizcarra, Zhi Li +2

Despite recent progress in video large language models (VideoLLMs), a key open challenge remains: how to equip models with chain-of-thought (CoT) reasoning abilities grounded in fi…

cs.LG2025

GTS-LUM: Reshaping User Behavior Modeling with LLMs in Telecommunications Industry

Liu Shi, Tianwu Zhou, Wei Xu +6

As telecommunication service providers shifting their focus to analyzing user behavior for package design and marketing interventions, a critical challenge lies in developing a uni…