11 papers
SCOPE-Router: Cost-Aware Open-Set VLM Routing for Execution-Oriented Tasks
Tao Yu, Yifei Qu, Zhiqing Cui +14
Model routing aims to select the most suitable model from a candidate pool for each query, balancing quality and cost. Existing VLM routing research is limited to traditional VQA e…
Diffusion Models are Open-World Affordance Learners: Leveraging Generative Priors for 3D Affordance Learning
Hanqing Wang, Zhenhao Zhang, Kaiyang Ji +12
3D affordance grounding aims to understand how diverse objects can be manipulated, making it a cornerstone of embodied interaction. However, prior works struggle to generalize to o…
Affordance-R1: Reinforcement Learning for Generalizable Affordance Reasoning in Multimodal Large Language Model
Hanqing Wang, Shaoyang Wang, Yiming Zhong +7
Affordance grounding focuses on predicting the specific regions of objects that are associated with the actions to be performed by robots. It plays a vital role in the fields of hu…
Omni-DeepSearch: A Benchmark for Audio-Driven Omni-Modal Deep Search
Tao Yu, yiming ding, Shenghua Chai +16
Current omni-modal benchmarks mainly evaluate models under settings where multiple modalities are provided simultaneously, while the ability to start from audio alone and actively…
Uno-Orchestra: Parsimonious Agent Routing via Selective Delegation
Zhiqing Cui, Haotong Xie, Jiahao Yuan +11
Large language model (LLM) multi-agent systems typically rely on rigid orchestration, committing either to flat per-query routing or to hand-engineered task decomposition, so decom…
PaperX: A Unified Framework for Multimodal Academic Presentation Generation with Scholar DAG
Tao Yu, Minghui Zhang, Zhiqing Cui +17
Transforming scientific papers into multimodal presentation content is essential for research dissemination but remains labor intensive. Existing automated solutions typically trea…