activity
20242026
collaborators

8 papers

cs.AI2026

Retrieval-Augmented Robots via Retrieve-Reason-Act

Izat Temiraliev, Diji Yang, Yi Zhang

To achieve general-purpose utility, we argue that robots must evolve from passive executors into active Information Retrieval users. In strictly zero-shot settings where no prior d…

cs.CV2026

SSR: Pushing the Limit of Spatial Intelligence with Structured Scene Reasoning

Yi Zhang, Youya Xia, Yong Wang +7

While Multimodal Large Language Models (MLLMs) excel in semantic tasks, they frequently lack the "spatial sense" essential for sophisticated geometric reasoning. Current models typ…

cs.AI2026

Worse than Zero-shot? A Fact-Checking Dataset for Evaluating the Robustness of RAG Against Misleading Retrievals

Linda Zeng, Rithwik Gupta, Divij Motwani +2

Retrieval-augmented generation (RAG) has shown impressive capabilities in mitigating hallucinations in large language models (LLMs). However, LLMs struggle to maintain consistent r…

cs.LG2026

Many Minds from One Model: Bayesian-Inspired Transformers for Population Diversity

Diji Yang, Yi Zhang

Despite their scale and success, modern transformers are usually trained as single-minded systems: optimization produces a deterministic set of parameters, representing a single fu…

cs.LG2025

Beyond Introspection: Reinforcing Thinking via Externalist Behavioral Feedback

Diji Yang, Linda Zeng, Kezhen Chen +1

While inference-time thinking allows Large Language Models (LLMs) to address complex problems, the extended thinking process can be unreliable or inconsistent because of the model'…

cs.CV2025

GenIR: Generative Visual Feedback for Mental Image Retrieval

Diji Yang, Minghao Liu, Chung-Hsiang Lo +2

Vision-language models (VLMs) have shown strong performance on text-to-image retrieval benchmarks. However, bridging this success to real-world applications remains a challenge. In…