7 papers
A Comprehensive Survey of Knowledge-Based Vision Question Answering Systems: The Lifecycle of Knowledge in Visual Reasoning Task
Jiaqi Deng, Zonghan Wu, Huan Huo +1
Knowledge-based Vision Question Answering (KB-VQA) extends general Vision Question Answering (VQA) by not only requiring the understanding of visual and textual inputs but also ext…
DART: Semantic Recoverability for Structured Tool Agents
Ke Yang, Panpan Li, Zonghan Wu +3
When a structured tool agent fails mid-execution, the runtime faces a dilemma: replaying the entire task is safe but wasteful, while restoring from a local checkpoint is efficient…
Rethinking Point Clouds as Sequences: A Causal Next-Token Predictive Learning Framework
Yumeng Yao, Jingzhi Dong, Haowen Gu +4
With the rapid progress of multimodal foundation models and predictive pre-training, an important open question is how to equip 3D point clouds with a pre-training paradigm that is…
MIRL: Mutual Information-Guided Reinforcement Learning for Vision-Language Models
Yin Zhang, Jiaxuan Zhao, Zonghan Wu +5
Vision-Language Models (VLMs) frequently suffer from visual perception errors and hallucinations that compromise answer accuracy in complex reasoning tasks. Reinforcement Learning…
PanFoMa: A Lightweight Foundation Model and Benchmark for Pan-Cancer
Xiaoshui Huang, Tianlin Zhu, Yifan Zuo +7
Single-cell RNA sequencing (scRNA-seq) is essential for decoding tumor heterogeneity. However, pan-cancer research still faces two key challenges: learning discriminative and effic…
Towards Competent AI for Fundamental Analysis in Finance: A Benchmark Dataset and Evaluation
Zonghan Wu, Congyuan Zou, Junlin Wang +3
Generative AI, particularly large language models (LLMs), is beginning to transform the financial industry by automating tasks and helping to make sense of complex financial inform…