collaborators

10 papers

cs.CL2026

Beyond Position Bias: Shifting Context Compression from Position-Driven to Semantic-Driven

Jiwei Tang, Zhijing Huang, Xinyu Zhang +5

Large Language Models (LLMs) have demonstrated exceptional performance across diverse tasks. However, their deployment in long-context scenarios faces high computational overhead a…

cs.IR2026

SAKE: Self-aware Knowledge Exploitation-Exploration for Grounded Multimodal Named Entity Recognition

Jielong Tang, Xujie Yuan, Jiayang Liu +6

Grounded Multimodal Named Entity Recognition (GMNER) aims to extract named entities and localize their visual regions within image-text pairs, serving as a pivotal capability for v…

cs.CV2025

An Iteration-Free Fixed-Point Estimator for Diffusion Inversion

Yifei Chen, Kaiyu Song, Yan Pan +3

Diffusion inversion aims to recover the initial noise corresponding to a given image such that this noise can reconstruct the original image through the denoising diffusion process…

cs.SD2025

HPSU: A Benchmark for Human-Level Perception in Real-World Spoken Speech Understanding

Chen Li, Peiji Yang, Yicheng Zhong +5

Recent advances in Speech Large Language Models (Speech LLMs) have led to great progress in speech understanding tasks such as Automatic Speech Recognition (ASR) and Speech Emotion…

cs.AI2025

CoS: Towards Optimal Event Scheduling via Chain-of-Scheduling

Yiming Zhao, Jiwei Tang, Shimin Di +3

Recommending event schedules is a key issue in Event-based Social Networks (EBSNs) in order to maintain user activity. An effective recommendation is required to maximize the user'…

cs.IR2025

ReFineG: Synergizing Small Supervised Models and LLMs for Low-Resource Grounded Multimodal NER

Jielong Tang, Shuang Wang, Zhenxing Wang +2

Grounded Multimodal Named Entity Recognition (GMNER) extends traditional NER by jointly detecting textual mentions and grounding them to visual regions. While existing supervised m…