5 papers
An Empirical Study of Validating Synthetic Data for Text-Based Person Retrieval
Min Cao, Yuxin Lu, Ziyin Zeng +3
Data plays a pivotal role in Text-Based Person Retrieval (TBPR) research. Mainstream research paradigm necessitates real-world person images with manual textual annotations for tra…
MME-Industry: A Cross-Industry Multimodal Evaluation Benchmark
Dongyi Yi, Guibo Zhu, Chenglin Ding +3
With the rapid advancement of Multimodal Large Language Models (MLLMs), numerous evaluation benchmarks have emerged. However, comprehensive assessments of their performance across…
AnyDesign: Versatile Area Fashion Editing via Mask-Free Diffusion
Yunfang Niu, Dong Yi, Lingxiang Wu +2
Fashion image editing aims to modify a person's appearance based on a given instruction. Existing methods require auxiliary tools like segmenters and keypoint extractors, lacking a…
Recurrent Context Compression: Efficiently Expanding the Context Window of LLM
Chensen Huang, Guibo Zhu, Xuepeng Wang +5
To extend the context length of Transformer-based large language models (LLMs) and improve comprehension capabilities, we often face limitations due to computational resources and…
PFDM: Parser-Free Virtual Try-on via Diffusion Model
Yunfang Niu, Dong Yi, Lingxiang Wu +3
Virtual try-on can significantly improve the garment shopping experiences in both online and in-store scenarios, attracting broad interest in computer vision. However, to achieve h…