11 papers
SCOPE-Router: Cost-Aware Open-Set VLM Routing for Execution-Oriented Tasks
Tao Yu, Yifei Qu, Zhiqing Cui +14
Model routing aims to select the most suitable model from a candidate pool for each query, balancing quality and cost. Existing VLM routing research is limited to traditional VQA e…
Omni-DeepSearch: A Benchmark for Audio-Driven Omni-Modal Deep Search
Tao Yu, yiming ding, Shenghua Chai +16
Current omni-modal benchmarks mainly evaluate models under settings where multiple modalities are provided simultaneously, while the ability to start from audio alone and actively…
Uno-Orchestra: Parsimonious Agent Routing via Selective Delegation
Zhiqing Cui, Haotong Xie, Jiahao Yuan +11
Large language model (LLM) multi-agent systems typically rely on rigid orchestration, committing either to flat per-query routing or to hand-engineered task decomposition, so decom…
PaperX: A Unified Framework for Multimodal Academic Presentation Generation with Scholar DAG
Tao Yu, Minghui Zhang, Zhiqing Cui +17
Transforming scientific papers into multimodal presentation content is essential for research dissemination but remains labor intensive. Existing automated solutions typically trea…
VideoAfford: Grounding 3D Affordance from Human-Object-Interaction Videos via Multimodal Large Language Model
Hanqing Wang, Mingyu Liu, Xiaoyu Chen +9
3D affordance grounding aims to highlight the actionable regions on 3D objects, which is crucial for robotic manipulation. Previous research primarily focused on learning affordanc…
Kardia-R1: Unleashing LLMs to Reason toward Understanding and Empathy for Emotional Support via Rubric-as-Judge Reinforcement Learning
Jiahao Yuan, Zhiqing Cui, Hanqing Wang +3
As web platforms evolve towards greater personalization and emotional complexity, conversational agents must transcend superficial empathy to demonstrate identity-aware emotional r…