collaborators

5 papers

cs.CV2025

COOPER: A Unified Model for Cooperative Perception and Reasoning in Spatial Intelligence

Zefeng Zhang, Xiangzhao Hao, Hengzhu Tang +8

Visual Spatial Reasoning is crucial for enabling Multimodal Large Language Models (MLLMs) to understand object properties and spatial relationships, yet current models still strugg…

cs.CL2025

Leveraging Generative Models for Real-Time Query-Driven Text Summarization in Large-Scale Web Search

Zeyu Xiong, Yixuan Nan, Li Gao +4

In the dynamic landscape of large-scale web search, Query-Driven Text Summarization (QDTS) aims to generate concise and informative summaries from textual documents based on a give…

cs.CV2025

Debiasing Multimodal Large Language Models via Noise-Aware Preference Optimization

Zefeng Zhang, Hengzhu Tang, Jiawei Sheng +6

Multimodal Large Language Models excel in various tasks, yet often struggle with modality bias, where the model tends to rely heavily on a single modality and overlook critical inf…

cs.CV2025

Multi-Branch Collaborative Learning Network for Video Quality Assessment in Industrial Video Search

Hengzhu Tang, Zefeng Zhang, Zhiping Li +5

Video Quality Assessment (VQA) is vital for large-scale video retrieval systems, aimed at identifying quality issues to prioritize high-quality videos. In industrial systems, low-q…

cs.IR2025

Exploring Preference-Guided Diffusion Model for Cross-Domain Recommendation

Xiaodong Li, Hengzhu Tang, Jiawei Sheng +5

Cross-domain recommendation (CDR) has been proven as a promising way to alleviate the cold-start issue, in which the most critical problem is how to draw an informative user repres…