collaborators

6 papers

cs.AI2026

MIRROR: Learning from the Other View for Multi-Modal Reasoning

Wen Ye, Yuxiao Qu, Aviral Kumar +1

Unlike large language models (LLMs) that exhibit strong reasoning capabilities, vision-language models (VLMs) struggle with visual reasoning, even on geometry problems that admit e…

cs.CV2026

An Exam for Active Observers

Jiarui Zhang, Muzi Tao, Shangshang Wang +3

Human vision is a closed loop: gaze is continuously redirected by intermediate hypotheses rather than a single snapshot. Decades of psychophysics and cognitive science have argued…

cs.CL2026

From Reasoning to Pixels: Benchmarking the Alignment Gap in Unified Multimodal Models

Cheng Yang, Chufan Shi, Bo Shui +7

Unified multimodal models (UMMs) aim to integrate multimodal understanding and generation within a unified architecture, yet it remains unclear to what extent their representations…

cs.CL2026

Don't Let It Hallucinate: Premise Verification via Retrieval-Augmented Logical Reasoning

Yuehan Qin, Shawn Li, Yi Nian +3

Large language models (LLMs) have shown substantial capacity for generating fluent, contextually appropriate responses. However, they can produce hallucinated outputs, especially w…

cs.CL2025

LLM The Genius Paradox: A Linguistic and Math Expert's Struggle with Simple Word-based Counting Problems

Nan Xu, Xuezhe Ma

Interestingly, LLMs yet struggle with some basic tasks that humans find trivial to handle, e.g., counting the number of character r's in the word "strawberry". There are several po…

cs.CL2025

DecoPrompt : Decoding Prompts Reduces Hallucinations when Large Language Models Meet False Premises

Nan Xu, Xuezhe Ma

While large language models (LLMs) have demonstrated increasing power, they have also called upon studies on their hallucinated outputs that deviate from factually correct statemen…