activity
20242026
most citedStreamSense: Streaming Social Task Detection with Selective Vision-Language Model Routing

1 citations · 1 across the 4 of their papers we have counts for

collaborators

7 papers

cs.CV20261 cited

StreamSense: Streaming Social Task Detection with Selective Vision-Language Model Routing

Han Wang, Deyi Ji, Lanyun Zhu +2

Live streaming platforms require real-time monitoring and reaction to social signals, utilizing partial and asynchronous evidence from video, text, and audio. We propose StreamSens…

cs.CV2025

Video-LMM Post-Training: A Deep Dive into Video Reasoning with Large Multimodal Models

Yolo Y. Tang, Jing Bi, Pinxin Liu +24

Video understanding represents the most challenging frontier in computer vision, requiring models to reason about complex spatiotemporal relationships, long-term dependencies, and…

cs.AI2025

Latent Chain-of-Thought for Visual Reasoning

Guohao Sun, Hang Hua, Jian Wang +5

Chain-of-thought (CoT) reasoning is critical for improving the interpretability and reliability of Large Vision-Language Models (LVLMs). However, existing training algorithms such…

cs.CV2025

Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal Prompting

Yunlong Tang, Jing Bi, Chao Huang +16

We present CAT-V (Caption AnyThing in Video), a training-free framework for fine-grained object-centric video captioning that enables detailed descriptions of user-selected objects…

cs.CV2024

SafeCFG: Controlling Harmful Features with Dynamic Safe Guidance for Safe Generation

Jiadong Pan, Liang Li, Hongcheng Gao +3

Diffusion models (DMs) have demonstrated exceptional performance in text-to-image tasks, leading to their widespread use. With the introduction of classifier-free guidance (CFG), t…

cs.CV2024

FINECAPTION: Compositional Image Captioning Focusing on Wherever You Want at Any Granularity

Hang Hua, Qing Liu, Lingzhi Zhang +5

The advent of large Vision-Language Models (VLMs) has significantly advanced multimodal tasks, enabling more sophisticated and accurate reasoning across various applications, inclu…