collaborators

6 papers

cs.AI2026

Grounded Chess Reasoning in Language Models via Master Distillation

Zhenwei Tang, Qianfeng Wen, Seth Grief-Albert +4

Language models often lack grounded reasoning capabilities in specialized domains where training data is scarce but bespoke systems excel. We introduce a general framework for dist…

cs.IR2026

SafeGEO: Understanding Generative Engine Optimization Risks in Recommendation Agents

Qianfeng Wen, Yifan Simon Liu, Xin Liu +4

Generative Engine Optimization (GEO) lets content owners rewrite web content to increase their visibility in generative systems. In recommendation agents, this creates a risk that…

cs.AI2026

ThinkTwice: Jointly Optimizing Large Language Models for Reasoning and Self-Refinement

Difan Jiao, Qianfeng Wen, Blair Yang +2

We introduce ThinkTwice, a simple two-phase framework that jointly optimizes LLMs to solve reasoning problems and refine the answers, based on Group Relative Policy Optimization (G…

cs.CL2026

OasisSimp: An Open-source Asian-English Sentence Simplification Dataset

Hannah Liu, Muxin Tian, Iqra Ali +8

Sentence simplification aims to make complex text more accessible by reducing linguistic complexity while preserving the original meaning. However, progress in this area remains li…

cs.SE2026

SWE-Bench Mobile: Can Large Language Model Agents Develop Industry-Level Mobile Applications?

Muxin Tian, Zhe Wang, Blair Yang +7

Can large language model agents develop industry-level mobile applications? We introduce \textbf{SWE-Bench Mobile}, a benchmark for evaluating coding agents on realistic software e…

cs.AI2025

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models

Zhenwei Tang, Difan Jiao, Blair Yang +1

Evaluating whether vision-language models (VLMs) reason consistently across representations is challenging because modality comparisons are typically confounded by task differences…