8 papers
UniSteer: Text-Guided Flow Matching in Activation Space for Versatile LLM Steering
Yingdong Shi, Ruiming Zhang, Changming Li +4
Activation-based control steers large language models (LLMs) by intervening on their internal representations during inference, and has emerged as an effective paradigm for control…
LFS: Learnable Frame Selector for Event-Aware and Temporally Diverse Video Captioning
Lianying Chao, Linfeng Yin, Peiyu Ren +8
Video captioning models convert frames into visual tokens and generate descriptions with large language models (LLMs). Since encoding all frames is prohibitively expensive, uniform…
Multi-Modal LLM based Image Captioning in ICT: Bridging the Gap Between General and Industry Domain
Lianying Chao, Kai Zhang, Haoran Cai +3
In the information and communications technology (ICT) industry, training a domain-specific large language model (LLM) or constructing a retrieval-augmented generation system requi…
Decoding the Delta: Unifying Remote Sensing Change Detection and Understanding with Multimodal Large Language Models
Xiaohe Li, Jiahao Li, Kaixin Zhang +5
While Multimodal Large Language Models (MLLMs) excel in general vision-language tasks, their application to remote sensing change understanding is hindered by a fundamental "tempor…
LADR: Locality-Aware Dynamic Rescue for Efficient Text-to-Image Generation with Diffusion Large Language Models
Chenglin Wang, Yucheng Zhou, Shawn Chen +2
Discrete Diffusion Language Models have emerged as a compelling paradigm for unified multimodal generation, yet their deployment is hindered by high inference latency arising from…
Clue Matters: Leveraging Latent Visual Clues to Empower Video Reasoning
Kaixin zhang, Xiaohe Li, Jiahao Li +4
Multi-modal Large Language Models (MLLMs) have significantly advanced video reasoning, yet Video Question Answering (VideoQA) remains challenging due to its demand for temporal cau…