collaborators

5 papers

cs.CV2025

Cross-Modal Causal Intervention for Medical Report Generation

Weixing Chen, Yang Liu, Ce Wang +4

Radiology Report Generation (RRG) is essential for computer-aided diagnosis and medication guidance, which can relieve the heavy burden of radiologists by automatically generating…

cs.CV2025

DreamArtist++: Controllable One-Shot Text-to-Image Generation via Positive-Negative Adapter

Ziyi Dong, Pengxu Wei, Liang Lin

State-of-the-arts text-to-image generation models such as Imagen and Stable Diffusion Model have succeed remarkable progresses in synthesizing high-quality, feature-rich images wit…

cs.CV2024

MEIA: Multimodal Embodied Perception and Interaction in Unknown Environments

Yang Liu, Xinshuai Song, Kaixuan Jiang +4

With the surge in the development of large language models, embodied intelligence has attracted increasing attention. Nevertheless, prior works on embodied intelligence typically e…

cs.CV2024

Heterogeneous Semantic Transfer for Multi-label Recognition with Partial Labels

Tianshui Chen, Tao Pu, Lingbo Liu +3

Multi-label image recognition with partial labels (MLR-PL), in which some labels are known while others are unknown for each image, may greatly reduce the cost of annotation and th…

cs.CV2024

TIP-Editor: An Accurate 3D Editor Following Both Text-Prompts And Image-Prompts

Jingyu Zhuang, Di Kang, Yan-Pei Cao +3

Text-driven 3D scene editing has gained significant attention owing to its convenience and user-friendliness. However, existing methods still lack accurate control of the specified…