collaborators

7 papers

cs.AI2026

SkillTrace: Traversing a Query-Skill Graph for Composable LLM Agents

Yue Yao, Shengyuan Wang, Xin Chen +4

Large language model agents increasingly solve complex tasks by composing reusable skills from a library. To address this, the key challenge is not merely to retrieve individually…

cs.CL2026

Adversarial Style Optimization: Enhancing VLM Jailbreaks by GRPO-based Stylistic Triggers Optimization

Bingjun Luo, Jialin Guo, Yue Yao +1

Multimodal Large Language Models (MLLMs) have achieved impressive performance, but their safety alignment remains vulnerable to jailbreak attacks. Existing content-based jailbreaks…

cs.AI2026

ST-SimDiff: Balancing Spatiotemporal Similarity and Difference for Efficient Video Understanding with MLLMs

Bingjun Luo, Tony Wang, Chaoqi Chen +1

Multimodal Large Language Models (MLLMs) face significant computational overhead when processing long videos due to the massive number of visual tokens required. To improve efficie…

cs.AI2026

Enhancing Visual Token Representations for Video Large Language Models via Training-Free Spatial-Temporal Pooling and Gridding

Bingjun Luo, Tony Wang, Hanqi Chen +1

Recent advances in Multimodal Large Language Models (MLLMs) have significantly advanced video understanding tasks, yet challenges remain in efficiently compressing visual tokens wh…

cs.CV2025

HunyuanOCR Technical Report

Hunyuan Vision Team, Pengyuan Lyu, Xingyu Wan +29

This paper presents HunyuanOCR, a commercial-grade, open-source, and lightweight (1B parameters) Vision-Language Model (VLM) dedicated to OCR tasks. The architecture comprises a Na…

cs.CV2025

Unleashing the Intrinsic Visual Representation Capability of Multimodal Large Language Models

Hengzhuang Li, Xinsong Zhang, Qiming Peng +6

Multimodal Large Language Models (MLLMs) have demonstrated remarkable proficiency in multimodal tasks. Despite their impressive performance, MLLMs suffer from the modality imbalanc…