activity
20242026
most citedVisual Prompting in Multimodal Large Language Models: A Survey

4 citations · 6 across the 6 of their papers we have counts for

collaborators

10 papers

cs.CV2026

SceneAlign: Aligning Multimodal Reasoning to Scene Graphs in Complex Visual Scenes

Chuhan Wang, Xintong Li, Jennifer Yuntong Zhang +5

Multimodal large language models often struggle with faithful reasoning in complex visual scenes, where intricate entities and relations require precise visual grounding at each st…

cs.IR2025

The 2nd Workshop on Human-Centered Recommender Systems

Kaike Zhang, Jiakai Tang, Du Su +6

Recommender systems shape how people discover information, form opinions, and connect with society. Yet, as their influence grows, traditional metrics, e.g., accuracy, clicks, and…

cs.IR2025

Listwise Preference Diffusion Optimization for User Behavior Trajectories Prediction

Hongtao Huang, Chengkai Huang, Junda Wu +3

Forecasting multi-step user behavior trajectories requires reasoning over structured preferences across future actions, a challenge overlooked by traditional sequential recommendat…

cs.CL2025

Pluralistic Off-policy Evaluation and Alignment

Chengkai Huang, Junda Wu, Zhouhang Xie +6

Personalized preference alignment for LLMs with diverse human preferences requires evaluation and alignment methods that capture pluralism. Most existing preference alignment datas…

cs.CL2025

SAND: Boosting LLM Agents with Self-Taught Action Deliberation

Yu Xia, Yiran Shen, Junda Wu +5

Large Language Model (LLM) agents are commonly tuned with supervised finetuning on ReAct-style expert trajectories or preference optimization over pairwise rollouts. Most of these…

cs.AI2025

DICE: Dynamic In-Context Example Selection in LLM Agents via Efficient Knowledge Transfer

Ruoyu Wang, Junda Wu, Yu Xia +4

Large language model-based agents, empowered by in-context learning (ICL), have demonstrated strong capabilities in complex reasoning and tool-use tasks. However, existing works ha…