7 papers
Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization
Wenhao Yang, Yu Xia, Jinlong Huang +10
Recent advancements in Multimodal Large Language Models (MLLMs) have incentivized models to ``think with images'' by actively invoking visual tools during multi-turn reasoning. The…
Distributed Online Convex Optimization with Efficient Communication: Improved Algorithm and Lower bounds
Sifan Yang, Wenhao Yang, Wei Jiang +1
We investigate distributed online convex optimization with compressed communication, where learners connected by a network collaboratively minimize a sequence of global loss fu…
Deep But Reliable: Advancing Multi-turn Reasoning for Thinking with Images
Wenhao Yang, Yu Xia, Jinlong Huang +7
Recent advances in large Vision-Language Models (VLMs) have exhibited strong reasoning capabilities on complex visual tasks by thinking with images in their Chain-of-Thought (CoT),…
Dual Adaptivity: Universal Algorithms for Minimizing the Adaptive Regret of Convex Functions
Lijun Zhang, Wenhao Yang, Guanghui Wang +2
To deal with changing environments, a new performance measure -- adaptive regret, defined as the maximum static regret over any interval, was proposed in online learning. Under the…
Improved Analysis for Sign-based Methods with Momentum Updates
Wei Jiang, Dingzhi Yu, Sifan Yang +2
In this paper, we present enhanced analysis for sign-based optimization algorithms with momentum updates. Traditional sign-based methods, under the separable smoothness assumption,…
Discounted Online Convex Optimization: Uniform Regret Across a Continuous Interval
Wenhao Yang, Sifan Yang, Lijun Zhang
Reflecting the greater significance of recent history over the distant past in non-stationary environments, -discounted regret has been introduced in online convex optimization…