activity
20242026
most citedLinVT: Empower Your Image-level Large Language Model to Understand Videos

1 citations · 1 across the 8 of their papers we have counts for

collaborators

8 papers

cs.CV2026

CORE-Seg: Reasoning-Driven Segmentation for Complex Lesions via Reinforcement Learning

Yuxin Xie, Yuming Chen, Yishan Yang +5

Medical image segmentation is undergoing a paradigm shift from conventional visual pattern matching to cognitive reasoning analysis. Although Multimodal Large Language Models (MLLM…

cs.CV2025

GISE-TTT:A Framework for Global InformationSegmentation and Enhancement

Fenglei Hao, Yuliang Yang, Ruiyuan Su +3

This paper addresses the challenge of capturing global temporaldependencies in long video sequences for Video Object Segmentation (VOS). Existing architectures often fail to effect…

cs.CV2025

HiMix: Reducing Computational Complexity in Large Vision-Language Models

Xuange Zhang, Dengjie Li, Bo Liu +7

Benefiting from recent advancements in large language models and modality alignment techniques, existing Large Vision-Language Models(LVLMs) have achieved prominent performance acr…

cs.CV2024

Manga Generation via Layout-controllable Diffusion

Siyu Chen, Dengjie Li, Zenghao Bao +4

Generating comics through text is widely studied. However, there are few studies on generating multi-panel Manga (Japanese comics) solely based on plain text. Japanese manga contai…

cs.CV2024★ 1 cited

LinVT: Empower Your Image-level Large Language Model to Understand Videos

Lishuai Gao, Yujie Zhong, Yingsen Zeng +3

Large Language Models (LLMs) have been widely used in various tasks, motivating us to develop an LLM-based assistant for videos. Instead of training from scratch, we propose a modu…

cs.CV2024

RFSR: Improving ISR Diffusion Models via Reward Feedback Learning

Xiaopeng Sun, Qinwei Lin, Yu Gao +6

Generative diffusion models (DM) have been extensively utilized in image super-resolution (ISR). Most of the existing methods adopt the denoising loss from DDPMs for model optimiza…