8 papers
MarkovScale: Towards Optimal Sequential Scaling at Inference Time
Youkang Wang, Jian Wang, Rubing Chen +3
Sequential scaling is a prominent inference-time scaling paradigm, yet its performance improvements are typically modest and not well understood, largely due to the prevalence of h…
OptPO: Optimal Rollout Allocation for Test-time Policy Optimization
Youkang Wang, Jian Wang, Rubing Chen +3
Test-time policy optimization enables large language models (LLMs) to adapt to distribution shifts by leveraging feedback from self-generated rollouts. However, existing methods re…
OptScale: Probabilistic Optimality for Inference-time Scaling
Youkang Wang, Jian Wang, Rubing Chen +1
Inference-time scaling has emerged as a powerful technique for enhancing the reasoning performance of Large Language Models (LLMs). However, existing approaches often rely on heuri…
Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning
Dayong Liang, Changmeng Zheng, Zhiyuan Wen +3
Traditional scene graphs primarily focus on spatial relationships, limiting vision-language models' (VLMs) ability to reason about complex interactions in visual scenes. This paper…
Cardiac Evidence Backtracking for Eating Behavior Monitoring using Collocative Electrocardiogram Imagining
Xu-Lu Zhang, Zhen-Qun Yang, Dong-Mei Jiang +4
Eating monitoring has remained an open challenge in medical research for years due to the lack of non-invasive sensors for continuous monitoring and the reliable methods for automa…
PolySmart @ TRECVid 2024 Video Captioning (VTT)
Jiaxin Wu, Wengyu Zhang, Xiao-Yong Wei +1
In this paper, we present our methods and results for the Video-To-Text (VTT) task at TRECVid 2024, exploring the capabilities of Vision-Language Models (VLMs) like LLaVA and LLaVA…