10 papers · 1 filter
VLNVerse: A Benchmark for Vision-Language Navigation with Versatile, Embodied, Realistic Simulation and Evaluation
Sihao Lin, Zerui Li, Xunyi Zhao +10
Despite remarkable progress in Vision-Language Navigation (VLN), existing benchmarks remain confined to fixed, small-scale datasets with naive physical simulation. These shortcomin…
WEAVE: Unleashing and Benchmarking the In-context Interleaved Comprehension and Generation
Wei Chow, Jiachun Pan, Yongyuan Liang +10
Recent advances in unified multimodal models (UMMs) have enabled impressive progress in visual comprehension and generation. However, existing datasets and benchmarks focus primari…
TAlignDiff: Automatic Tooth Alignment assisted by Diffusion-based Transformation Learning
Yunbi Liu, Enqi Tang, Shiyu Li +5
Orthodontic treatment hinges on tooth alignment, which significantly affects occlusal function, facial aesthetics, and patients' quality of life. Current deep learning approaches p…
On Path to Multimodal Generalist: General-Level and General-Bench
Hao Fei, Yuan Zhou, Juncheng Li +29
The Multimodal Large Language Model (MLLM) is currently experiencing rapid growth, driven by the advanced capabilities of LLMs. Unlike earlier specialists, existing MLLMs are evolv…
HumanEdit: A High-Quality Human-Rewarded Dataset for Instruction-based Image Editing
Jinbin Bai, Wei Chow, Ling Yang +4
We present HumanEdit, a high-quality, human-rewarded dataset specifically designed for instruction-guided image editing, enabling precise and diverse image manipulations through op…
RepText: Rendering Visual Text via Replicating
Haofan Wang, Yujia Xu, Yimeng Li +5
Although contemporary text-to-image generation models have achieved remarkable breakthroughs in producing visually appealing images, their capacity to generate precise and flexible…