collaborators

5 papers

cs.CV2026

StreamTTT: Reconciling Real-Time Perception and Long-Term Memory in Streaming VLMs

Joya Chen, Zeyun Zhong, Mike Zheng Shou

Humans effortlessly perceive the present while remembering the past, yet streaming VLMs often trade off real-time perception against long-term memory. Prior work shows that shorten…

cs.CL2026

What Transfers from Text to Vision? Capability Scaling Laws and Transfer Dynamics for VLMs

Ziran Li, Qiang Wang, Zhengyu Chen +4

Choosing the right large language model (LLM) backbone is the most consequential decision when building a vision-language model (VLM), yet it remains fundamentally unprincipled: co…

cs.AI2024

Multi-Dimensional Insights: Benchmarking Real-World Personalization in Large Multimodal Models

YiFan Zhang, Shanglin Lei, Runqi Qiao +10

The rapidly developing field of large multimodal models (LMMs) has led to the emergence of diverse models with remarkable capabilities. However, existing benchmarks fail to compreh…

cs.CL2024

InstructERC: Reforming Emotion Recognition in Conversation with Multi-task Retrieval-Augmented Large Language Models

Shanglin Lei, Guanting Dong, Xiaoping Wang +3

The field of emotion recognition of conversation (ERC) has been focusing on separating sentence feature encoding and context modeling, lacking exploration in generative paradigms b…

cs.AI2024

We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?

Runqi Qiao, Qiuna Tan, Guanting Dong +15

Visual mathematical reasoning, as a fundamental visual reasoning ability, has received widespread attention from the Large Multimodal Models (LMMs) community. Existing benchmarks,…