activity
20242026
collaborators

8 papers

cs.LG2026

MarkovScale: Towards Optimal Sequential Scaling at Inference Time

Youkang Wang, Jian Wang, Rubing Chen +3

Sequential scaling is a prominent inference-time scaling paradigm, yet its performance improvements are typically modest and not well understood, largely due to the prevalence of h…

cs.LG2025

OptPO: Optimal Rollout Allocation for Test-time Policy Optimization

Youkang Wang, Jian Wang, Rubing Chen +3

Test-time policy optimization enables large language models (LLMs) to adapt to distribution shifts by leveraging feedback from self-generated rollouts. However, existing methods re…

cs.LG2025

OptScale: Probabilistic Optimality for Inference-time Scaling

Youkang Wang, Jian Wang, Rubing Chen +1

Inference-time scaling has emerged as a powerful technique for enhancing the reasoning performance of Large Language Models (LLMs). However, existing approaches often rely on heuri…

cs.CV2025

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning

Dayong Liang, Changmeng Zheng, Zhiyuan Wen +3

Traditional scene graphs primarily focus on spatial relationships, limiting vision-language models' (VLMs) ability to reason about complex interactions in visual scenes. This paper…

cs.LG2025

Cardiac Evidence Backtracking for Eating Behavior Monitoring using Collocative Electrocardiogram Imagining

Xu-Lu Zhang, Zhen-Qun Yang, Dong-Mei Jiang +4

Eating monitoring has remained an open challenge in medical research for years due to the lack of non-invasive sensors for continuous monitoring and the reliable methods for automa…

cs.CV2025

PolySmart @ TRECVid 2024 Video Captioning (VTT)

Jiaxin Wu, Wengyu Zhang, Xiao-Yong Wei +1

In this paper, we present our methods and results for the Video-To-Text (VTT) task at TRECVid 2024, exploring the capabilities of Vision-Language Models (VLMs) like LLaVA and LLaVA…