10 papers
Teaching MLLMs to Say No: Generalized Referring Expression Comprehension via Refusal Calibrated GRPO
Xuzheng Yang, Jun Ling, Tao Huang +2
We tackle the challenging yet underexplored task of Generalized Referring Expression Comprehension (GREC), which requires a model to localize the object described by a textual expr…
GOTS: Greedy Orthogonal Token Selection for High-Resolution Vision-Language Models
Jun Ling, Tao Huang, Junzhuo Liu +2
Modern vision-language models (VLMs) increasingly rely on dynamic or high-resolution visual encoding, producing thousands of visual tokens that substantially increase downstream la…
MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE
Junzhe Li, Yutao Cui, Tao Huang +8
Although GRPO substantially enhances flow matching models in human preference alignment of image generation, methods such as FlowGRPO and DanceGRPO still exhibit inefficiency due t…
Spatial4D-Bench: A Versatile 4D Spatial Intelligence Benchmark
Pan Wang, Yang Liu, Guile Wu +23
4D spatial intelligence involves perceiving and processing how objects move or change over time. Humans naturally possess 4D spatial intelligence, supporting a broad spectrum of sp…
Structured Self-Consistency:A Multi-Task Evaluation of LLMs on VirtualHome
Jiaqi Xu, Tao Huang, Kai Zhang
Embodied AI requires agents to understand goals, plan actions, and execute tasks in simulated environments. We present a comprehensive evaluation of Large Language Models (LLMs) on…
Distilling Cross-Modal Knowledge via Feature Disentanglement
Junhong Liu, Yuan Zhang, Tao Huang +2
Knowledge distillation (KD) has proven highly effective for compressing large models and enhancing the performance of smaller ones. However, its effectiveness diminishes in cross-m…