collaborators
Showing cs.CVShow all

20 papers · 1 filter

cs.CV2025

ObjectAdd: Adding Objects into Image via a Training-Free Diffusion Modification Fashion

Ziyue Zhang, Mingbao Lin, Quanjian Song +2

We introduce ObjectAdd, a training-free diffusion modification method to add user-expected objects into user-specified area. The motive of ObjectAdd stems from: first, describing e…

cs.CV2025

Parallel Vision Token Scheduling for Fast and Accurate Multimodal LMMs Inference

Wengyi Zhan, Mingbao Lin, Zhihang Lin +1

Multimodal large language models (MLLMs) deliver impressive vision-language reasoning but suffer steep inference latency because self-attention scales quadratically with sequence l…

cs.CV2025

Test-Time Temporal Sampling for Efficient MLLM Video Understanding

Kaibin Wang, Mingbao Lin

Processing long videos with multimodal large language models (MLLMs) poses a significant computational challenge, as the model's self-attention mechanism scales quadratically with…

cs.CV2025

UniVST: A Unified Framework for Training-free Localized Video Style Transfer

Quanjian Song, Mingbao Lin, Wengyi Zhan +3

This paper presents UniVST, a unified framework for localized video style transfer based on diffusion models. It operates without the need for training, offering a distinct advanta…

cs.CV2025

I&S-ViT: An Inclusive & Stable Method for Pushing the Limit of Post-Training ViTs Quantization

Yunshan Zhong, Jiawei Hu, Mingbao lin +2

Albeit the scalable performance of vision transformers (ViTs), the dense computational costs (training & inference) undermine their position in industrial applications. Post-traini…

cs.CV2025

DSNet: Detail-Semantic Deep Supervision Network for Medical Image Segmentation

Zhaohong Huang, Yuxin Zhang, Taojian Zhou +2

Deep Supervision Networks exhibit significant efficacy for the medical imaging community. Nevertheless, existing work merely supervises either the coarse-grained semantic features…