most citedReasoning Beyond Chain-of-Thought: A Latent Computational Mode in Large Language Models

1 citations · 1 across the 6 of their papers we have counts for

collaborators

8 papers

cs.CV2026

Towards Robustness against Typographic Attack with Training-free Concept Localization

Bohan Liu, Wenqian Ye, Guangzhi Xiong +3

Models trained via Contrastive Language-Image Pretraining (CLIP) serve as the foundational vision encoders for most modern Large Vision Language Models (LVLMs). Despite their wides…

cs.CV2026

Attention-guided Fine-tuning of Multimodal Large Language Models Improves Chain-of-Thought Reasoning

Sanchit Sinha, Guangzhi Xiong, Bohan Liu +2

The effectiveness of Chain-of-Thought (CoT) prompting in Multimodal Large Language Models (MLLMs) remains uncertain: across several visual reasoning benchmarks, CoT prompting often…

cs.CV2026

Retrieving Counterfactuals Improves Visual In-Context Learning

Guangzhi Xiong, Sanchit Sinha, Zhenghao He +1

Vision-language models (VLMs) have achieved impressive performance across a wide range of multimodal reasoning tasks, but they often struggle to disentangle fine-grained visual att…

cs.CL2026

Toward Faithful Retrieval-Augmented Generation with Sparse Autoencoders

Guangzhi Xiong, Zhenghao He, Bohan Liu +2

Retrieval-Augmented Generation (RAG) improves the factuality of large language models (LLMs) by grounding outputs in retrieved evidence, but faithfulness failures, where generation…

cs.LG2026

CASL: Concept-Aligned Sparse Latents for Interpreting Diffusion Models

Zhenghao He, Guangzhi Xiong, Boyang Wang +2

Internal activations of diffusion models encode rich semantic information, but interpreting such representations remains challenging. While Sparse Autoencoders (SAEs) have shown pr…

cs.CL20261 cited

Reasoning Beyond Chain-of-Thought: A Latent Computational Mode in Large Language Models

Zhenghao He, Guangzhi Xiong, Bohan Liu +2

Chain-of-Thought (CoT) prompting has improved the reasoning performance of large language models (LLMs), but it remains unclear why it works and whether it is the unique mechanism…