8 papers
Latent Bias Alignment for High-Fidelity Diffusion Inversion in Real-World Image Reconstruction and Manipulation
Weiming Chen, Qifan Liu, Siyi Liu +4
Recent research has shown that text-to-image diffusion models are capable of generating high-quality images guided by text prompts. But can they be used to generate or approximate…
Unleashing Video Language Models for Fine-grained HRCT Report Generation
Yingying Fang, Huichi Zhou, KinHei Lee +4
Generating precise diagnostic reports from High-Resolution Computed Tomography (HRCT) is critical for clinical workflow, yet it remains a formidable challenge due to the high patho…
PaAgent: Portrait-Aware Image Restoration Agent via Subjective-Objective Reinforcement Learning
Yijian Wang, Qingsen Yan, Jiantao Zhou +2
Image Restoration (IR) agents, leveraging multimodal large language models to perceive degradation and invoke restoration tools, have shown promise in automating IR tasks. However,…
CVBench: Benchmarking Cross-Video Synergies for Complex Multimodal Reasoning
Nannan Zhu, Yonghao Dong, Teng Wang +9
While multimodal large language models (MLLMs) exhibit strong performance on single-video tasks (e.g., video question answering), their capability for spatiotemporal pattern reason…
Understanding the Implicit User Intention via Reasoning with Large Language Model for Image Editing
Yijia Wang, Yiqing Shen, Weiming Chen +1
Existing image editing methods can handle simple editing instructions very well. To deal with complex editing instructions, they often need to jointly fine-tune the large language…
Generative Semantic Coding for Ultra-Low Bitrate Visual Communication and Analysis
Weiming Chen, Yijia Wang, Zhihan Zhu +1
We consider the problem of ultra-low bit rate visual communication for remote vision analysis, human interactions and control in challenging scenarios with very low communication b…