papers

Publications (28)

cs.CV2025

Video-Thinker: Sparking "Thinking with Videos" via Reinforcement Learning

Shijian Wang, Jiarui Jin, Xingjian Wang +6

Recent advances in image reasoning methods, particularly "Thinking with Images", have demonstrated remarkable success in Multimodal Large Language Models (MLLMs); however, this dyn…

cs.CL2024

Better Explain Transformers by Illuminating Important Information

Linxin Song, Yan Cui, Ao Luo +2

Transformer-based models excel in various natural language processing (NLP) tasks, attracting countless efforts to explain their inner workings. Prior methods explain Transformers…

cs.LG2022

Leveraging Instance Features for Label Aggregation in Programmatic Weak Supervision

Jieyu Zhang, Linxin Song, Alexander Ratner

Programmatic Weak Supervision (PWS) has emerged as a widespread paradigm to synthesize training labels efficiently. The core component of PWS is the label model, which infers true…

cs.CL2022

Adaptive Ranking-based Sample Selection for Weakly Supervised Class-imbalanced Text Classification

Linxin Song, Jieyu Zhang, Tianxiang Yang +1

To obtain a large amount of training labels inexpensively, researchers have recently adopted the weak supervision (WS) paradigm, which leverages labeling rules to synthesize traini…

cs.CL2023

NLPBench: Evaluating Large Language Models on Solving NLP Problems

Linxin Song, Jieyu Zhang, Lechao Cheng +3

Recent developments in large language models (LLMs) have shown promise in enhancing the capabilities of natural language processing (NLP). Despite these successes, there remains a…

cs.CL2025

The Hallucination Tax of Reinforcement Finetuning

Linxin Song, Taiwei Shi, Jieyu Zhao

Reinforcement finetuning (RFT) has become a standard approach for enhancing the reasoning capabilities of large language models (LLMs). However, its impact on model trustworthiness…

cs.CV2024

ProVision: Programmatically Scaling Vision-centric Instruction Data for Multimodal Language Models

Jieyu Zhang, Le Xue, Linxin Song +11

With the rise of multimodal applications, instruction data has become critical for training multimodal language models capable of understanding complex image-based queries. Existin…

cs.CL2026

CoAct-1: Computer-using Multi-Agent System with Coding Actions

Linxin Song, Yutong Dai, Viraj Prabhu +9

Autonomous agents that operate computers via Graphical User Interfaces (GUIs) often struggle with efficiency and reliability on complex, long-horizon tasks. While augmenting these…

cs.CL2026

Memory Retrieval for Changing Preferences

Yuehan Qin, Li Li, Linxin Song +4

Long-context dialogue systems must decide both when to access memory and which parts of the interaction history are relevant. Existing approaches typically rely on heuristic retrie…

cs.CL2025

Adaptive In-conversation Team Building for Language Model Agents

Linxin Song, Jiale Liu, Jieyu Zhang +5

Leveraging multiple large language model (LLM) agents has shown to be a promising approach for tackling complex tasks, while the effective design of multiple agents for a particula…

cs.CL2025

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations

Li Li, Peilin Cai, Ryan A. Rossi +21

We present PersonaConvBench, a large-scale benchmark for evaluating personalized reasoning and generation in multi-turn conversations with large language models (LLMs). Unlike exis…

cs.CL2026

Mitigating Bias in Locally Constrained Decoding via Tractable Proposals

Meihua Dang, Linxin Song, Honghua Zhang +3

Generations from large language models often fail to conform to desired constraints such as JSON schema. Existing locally constrained decoding (LCD) approaches enforce constraints…

cs.CV2025

Investigating the Scaling Effect of Instruction Templates for Training Multimodal Language Model

Shijian Wang, Linxin Song, Jieyu Zhang +9

Current multimodal language model (MLM) training approaches overlook the influence of instruction templates. Previous research deals with this problem by leveraging hand-crafted or…

cs.AI2024

Offline Training of Language Model Agents with Functions as Learnable Weights

Shaokun Zhang, Jieyu Zhang, Jiale Liu +4

Researchers and practitioners have recently reframed powerful Large Language Models (LLMs) as agents, enabling them to automate complex tasks largely via the use of specialized fun…

cs.LG2023

Taming Small-sample Bias in Low-budget Active Learning

Linxin Song, Jieyu Zhang, Xiaotian Lu +1

Active learning (AL) aims to minimize the annotation cost by only querying a few informative examples for each model training stage. However, training a model on a few queried exam…

cs.CR2026

The Blind Spot of Agent Safety: How Benign User Instructions Expose Critical Vulnerabilities in Computer-Use Agents

Xuwei Ding, Skylar Zhai, Linxin Song +6

Computer-use agents (CUAs) can now autonomously complete complex tasks in real digital environments, but when misled, they can also be used to automate harmful actions programmatic…

cs.CL2025

Discovering Knowledge Deficiencies of Language Models on Massive Knowledge Base

Linxin Song, Xuwei Ding, Jieyu Zhang +6

Large language models (LLMs) possess impressive linguistic capabilities but often fail to faithfully retain factual knowledge, leading to hallucinations and unreliable outputs. Und…

cs.LG2026

The Strongest Teacher Is Not Always the Best Teacher: Student-Centric Answer Selection

Zhengyu Hu, Zheyuan Xiao, Linxin Song +10

LLM training increasingly relies on teacher-generated supervision, from synthetic responses to reasoning traces and tool-use demonstrations. Current practice often chooses the high…

cs.CV2026

Video-Based Reward Modeling for Computer-Use Agents

Linxin Song, Jieyu Zhang, Huanxin Sheng +6

Computer-using agents (CUAs) are becoming increasingly capable; however, it remains difficult to scale evaluation of whether a trajectory truly fulfills a user instruction. In this…

cs.LG2025

Disentangling Likes and Dislikes in Personalized Generative Explainable Recommendation

Ryotaro Shimizu, Takashi Wada, Yu Wang +9

Recent research on explainable recommendation generally frames the task as a standard text generation problem, and evaluates models simply based on the textual similarity between t…

cs.CV2024

LARE: Latent Augmentation using Regional Embedding with Vision-Language Model

Kosuke Sakurai, Tatsuya Ishii, Ryotaro Shimizu +2

In recent years, considerable research has been conducted on vision-language models that handle both image and text data; these models are being applied to diverse downstream tasks…

cs.CV2024

SCP: Spherical-Coordinate-based Learned Point Cloud Compression

Ao Luo, Linxin Song, Keisuke Nonaka +4

In recent years, the task of learned point cloud compression has gained prominence. An important type of point cloud, the spinning LiDAR point cloud, is generated by spinning LiDAR…

cs.LG2026

Efficient Reinforcement Finetuning via Adaptive Curriculum Learning

Taiwei Shi, Yiyang Wu, Linxin Song +2

Reinforcement finetuning (RFT) has shown great potential for enhancing the mathematical reasoning capabilities of large language models (LLMs), but it is often sample- and compute-…

cs.SE2026

Converted, Not Equivalent: Benchmarking Codebase Conversion via Observational Equivalence

Linxin Song, Jiefeng Chen, Yue Huang +5

Coding agents increasingly act as codebase-scale collaborators that can assist with codebase conversion, but this progress has exposed a critical weakness: agents often over-trust…

cs.CL2026

Self-Evolving LLM Memory Extraction Across Heterogeneous Tasks

Yuqing Yang, Tengxiao Liu, Wang Bill Zhu +3

As LLM-based assistants become persistent and personalized, they must extract and retain useful information from past conversations as memory. However, the types of information wor…

cs.LG2026

Experiential Reinforcement Learning

Taiwei Shi, Sihao Chen, Bowen Jiang +3

Reinforcement learning has become the central approach for language models (LMs) to learn from environmental reward or feedback. In practice, the environmental feedback is usually…

cs.LG2025

Explaining Length Bias in LLM-Based Preference Evaluations

Zhengyu Hu, Linxin Song, Jieyu Zhang +7

The use of large language models (LLMs) as judges, particularly in preference comparisons, has become widespread, but this reveals a notable bias towards longer responses, undermin…

cs.CV2025

Attributed Synthetic Data Generation for Zero-shot Domain-specific Image Classification

Shijian Wang, Linxin Song, Ryotaro Shimizu +2

Zero-shot domain-specific image classification is challenging in classifying real images without ground-truth in-domain training examples. Recent research involved knowledge from t…