3 papers
cs.AI2026
ItinBench: Benchmarking Planning Across Multiple Cognitive Dimensions with Large Language Models
Tianlong Wang, Pinqiao Wang, Weili Shi +1
Large language models (LLMs) with advanced cognitive capabilities are emerging as agents for various reasoning and planning tasks. Traditional evaluations often focus on specific r…
cs.CL2026
Finding the Cracks: Improving LLMs Reasoning with Paraphrastic Probing and Consistency Verification
Weili Shi, Dongliang Guo, Lehan Yang +3
Large language models have demonstrated impressive performance across a variety of reasoning tasks. However, their problem-solving ability often declines on more complex tasks due…
cs.CV2025
VRMDiff: Text-Guided Video Referring Matting Generation of Diffusion
Lehan Yang, Jincen Song, Tianlong Wang +4
We propose a new task, video referring matting, which obtains the alpha matte of a specified instance by inputting a referring caption. We treat the dense prediction task of mattin…