activity
20242026
most citedVisual Prompting in Multimodal Large Language Models: A Survey

4 citations · 10 across the 15 of their papers we have counts for

collaborators
Showing 2026Show all

5 papers · 1 filter

cs.AI2026

AUSO: Action-Level Unified Skill Optimization from Internalization to Utilization

Huizu Lin, Chengkai Huang, Tianqi Gao +5

Skills play different roles as an agent's policy evolves: they should first provide learnable knowledge, then support capability formation, and finally be invoked only when they im…

cs.AI2026

Can We Break LLMs Out of Self-Loops? Fine-Grained Reasoning Control with Activation Steering

Sheldon Yu, Tong Yu, Xunyi Jiang +6

Extended reasoning has become standard for frontier Large Language Models (LLMs), yet the trajectories these models produce remain largely uncontrollable. Existing methods for shap…

cs.CL2026

FERA: Uncertainty-Aware Federated Reasoning for Large Language Models

Ruhan Wang, Chengkai Huang, Zhiyong Wang +6

Large language models (LLMs) exhibit strong reasoning capabilities when guided by high-quality demonstrations, yet such data is often distributed across organizations that cannot c…

cs.LG2026

Skill-CMIB: Multimodal Agent Skill for Consistent Action via Conditional Multimodal Information Bottleneck

Zihan Huang, Junda Wu, Tong Yu +6

While LLM-based agents excel at planning and executing long action sequences, their execution often remains inconsistent across trials, limiting reliability. Consolidating agent co…

cs.CV2026

SceneAlign: Aligning Multimodal Reasoning to Scene Graphs in Complex Visual Scenes

Chuhan Wang, Xintong Li, Jennifer Yuntong Zhang +5

Multimodal large language models often struggle with faithful reasoning in complex visual scenes, where intricate entities and relations require precise visual grounding at each st…