collaborators

7 papers

cs.CL2026

TypedCSIP: Typed Counterfactual Pretraining for Chinese Legislative Conflict Classification

Yao Liu

TypedCSIP is a typed counterfactual pretraining method for the conflict-classification task of the LCR-CN benchmark (Zhao et al., 2026): given a (superior, subordinate) provision p…

cs.CV2026

DUEL: Adversarial Self-Play for Multimodal Reasoning

Lin Qiu, Hanqing Zeng, Yao Liu +3

Reinforcement learning (RL) has emerged as an effective paradigm for improving the reasoning capability of vision-language models (VLMs). However, RL-based optimization typically d…

cs.CV2026

Efficient Long-Context Modeling in Diffusion Language Models via Block Approximate Sparse Attention

Wenhu Zhang, Yiming Wu, Huanyu Wang +6

Diffusion Language Models (DLMs) enable globally coherent, bidirectional, and controllable text generation, offering advantages over traditional autoregressive LLMs, while scaling…

cs.LG2025

Offline Learning and Forgetting for Reasoning with Large Language Models

Tianwei Ni, Allen Nie, Sapana Chaudhary +3

Leveraging inference-time search in large language models has proven effective in further enhancing a trained model's capability to solve complex mathematical and reasoning problem…

cs.AI2025

AgentOccam: A Simple Yet Strong Baseline for LLM-Based Web Agents

Ke Yang, Yao Liu, Sapana Chaudhary +4

Autonomy via agents using large language models (LLMs) for personalized, standardized tasks boosts human efficiency. Automating web tasks (like booking hotels within a budget) is i…

cs.LG2025

From Demonstrations to Rewards: Alignment Without Explicit Human Preferences

Siliang Zeng, Yao Liu, Huzefa Rangwala +3

One of the challenges of aligning large models with human preferences lies in both the data requirements and the technical complexities of current approaches. Predominant methods,…