activity
20242026
collaborators

5 papers

cs.SD2026

Don't Just Listen, Try Planning: Graph-based Retrieval-Generation Agent for Long-form Audio Meeting Understanding

Quanwei Tang, Dong Zhang, Shoushan Li +1

While long-form audio meeting understanding (LAMU) is garnering growing attention, task-specific question answering (QA) datasets remain scarce. Existing speech QA paradigms and st…

cs.AI2026

Relative Time Intervals Representation for Word-level Timestamping with Masked Training

Quanwei Tang, Zhiyu Tang, Xu Li +3

Although Speech Large Language Models (SpeechLLMs) excel at speech understanding and generation, their capacity for fine-grained, temporally aligned outputs remains underexplored.…

cs.CL2025

Zero-shot Cross-lingual NER via Mitigating Language Difference: An Entity-aligned Translation Perspective

Zhihao Zhang, Sophia Yat Mei Lee, Dong Zhang +2

Cross-lingual Named Entity Recognition (CL-NER) aims to transfer knowledge from high-resource languages to low-resource languages. However, existing zero-shot CL-NER (ZCL-NER) appr…

cs.CL2025

A Comprehensive Graph Framework for Question Answering with Mode-Seeking Preference Alignment

Quanwei Tang, Sophia Yat Mei Lee, Junshuang Wu +4

Recent advancements in retrieval-augmented generation (RAG) have enhanced large language models in question answering by integrating external knowledge. However, challenges persist…

cs.CV2024

Comment-aided Video-Language Alignment via Contrastive Pre-training for Short-form Video Humor Detection

Yang Liu, Tongfei Shen, Dong Zhang +3

The growing importance of multi-modal humor detection within affective computing correlates with the expanding influence of short-form video sharing on social media platforms. In t…