1 citations · 1 across the 6 of their papers we have counts for
10 papers
OracleAgent: A Multimodal Reasoning Agent for Oracle Bone Script Research
Caoshuo Li, Zengmao Ding, Xiaobin Hu +13
As one of the earliest writing systems, Oracle Bone Script (OBS) preserves the cultural and intellectual heritage of ancient civilizations. However, current OBS research faces two…
Human-MME: A Holistic Evaluation Benchmark for Human-Centric Multimodal Large Language Models
Yuansen Liu, Haiming Tang, Jinlong Peng +12
Multimodal Large Language Models (MLLMs) have demonstrated significant advances in visual understanding tasks. However, their capacity to comprehend human-centric scenes has rarely…
StrandDesigner: Towards Practical Strand Generation with Sketch Guidance
Na Zhang, Moran Li, Chengming Xu +6
Realistic hair strand generation is crucial for applications like computer graphics and virtual reality. While diffusion models can generate hairstyles from text or images, these i…
OracleFusion: Assisting the Decipherment of Oracle Bone Script with Structurally Constrained Semantic Typography
Caoshuo Li, Zengmao Ding, Xiaobin Hu +10
As one of the earliest ancient languages, Oracle Bone Script (OBS) encapsulates the cultural records and intellectual expressions of ancient civilizations. Despite the discovery of…
VTBench: Comprehensive Benchmark Suite Towards Real-World Virtual Try-on Models
Hu Xiaobin, Liang Yujie, Luo Donghao +5
While virtual try-on has achieved significant progress, evaluating these models towards real-world scenarios remains a challenge. A comprehensive benchmark is essential for three k…
When Preferences Diverge: Aligning Diffusion Models with Minority-Aware Adaptive DPO
Lingfan Zhang, Chen Liu, Chengming Xu +5
In recent years, the field of image generation has witnessed significant advancements, particularly in fine-tuning methods that align models with universal human preferences. This…