papers

Publications (30)

cs.IR2026

Efficient Sequential Recommendation for Long Term User Interest Via Personalization

Qiang Zhang, Hanchao Yu, Ivan Ji +14

Recent years have witnessed success of sequential modeling, generative recommender, and large language model for recommendation. Though the scaling law has been validated for seque…

cs.LG2025

BRIDLE: Generalized Self-supervised Learning with Quantization

Hoang M. Nguyen, Satya N. Shukla, Qiang Zhang +5

Self-supervised learning has been a powerful approach for learning meaningful representations from unlabeled data across various domains, reducing the reliance on large labeled dat…

cs.LG2025

Learning Critically: Selective Self Distillation in Federated Learning on Non-IID Data

Yuting He, Yiqiang Chen, XiaoDong Yang +3

Federated learning (FL) enables multiple clients to collaboratively train a global model while keeping local data decentralized. Data heterogeneity (non-IID) across clients has imp…

cs.AI2026

GISTBench: Evaluating LLM User Understanding via Evidence-Based Interest Verification

Iordanis Fostiropoulos, Muhammad Rafay Azhar, Abdalaziz Sawwan +8

We introduce GISTBench, a benchmark for evaluating Large Language Models' (LLMs) ability to understand users from their interaction histories in recommendation systems. Unlike trad…

eess.IV2021

Study Group Learning: Improving Retinal Vessel Segmentation Trained with Noisy Labels

Yuqian Zhou, Hanchao Yu, Humphrey Shi

Retinal vessel segmentation from retinal images is an essential task for developing the computer-aided diagnosis system for retinal diseases. Efforts have been made on high-perform…

cs.LG2022

Uniform Masking Prevails in Vision-Language Pretraining

Siddharth Verma, Yuchen Lu, Rui Hou +4

Masked Language Modeling (MLM) has proven to be an essential component of Vision-Language (VL) pretraining. To implement MLM, the researcher must make two design choices: the maski…

cs.IR2026

Verifiable Reasoning for LLM-based Generative Recommendation

Xinyu Lin, Hanqing Zeng, Hanchao Yu +8

Reasoning in Large Language Models (LLMs) has recently shown strong potential in enhancing generative recommendation through deep understanding of complex user preference. Existing…

cs.CV2023

MMViT: Multiscale Multiview Vision Transformers

Yuchen Liu, Natasha Ong, Kaiyan Peng +8

We present Multiscale Multiview Vision Transformers (MMViT), which introduces multiscale feature maps and multiview encodings to transformer models. Our model encodes different vie…

eess.IV2020

Motion Pyramid Networks for Accurate and Efficient Cardiac Motion Estimation

Hanchao Yu, Xiao Chen, Humphrey Shi +3

Cardiac motion estimation plays a key role in MRI cardiac feature tracking and function assessment such as myocardium strain. In this paper, we propose Motion Pyramid Networks, a n…

cs.CL2023

RoAST: Robustifying Language Models via Adversarial Perturbation with Selective Training

Jaehyung Kim, Yuning Mao, Rui Hou +7

Fine-tuning pre-trained language models (LMs) has become the de facto standard in many NLP tasks. Nevertheless, fine-tuned LMs are still prone to robustness issues, such as adversa…

cs.CL2025

RESTRAIN: From Spurious Votes to Signals -- Self-Driven RL with Self-Penalization

Zhaoning Yu, Will Su, Leitian Tao +9

Reinforcement learning with human-annotated data has boosted chain-of-thought reasoning in large reasoning models, but these gains come at high costs in labeled data while falterin…

cs.IR2026

Don't Waste It: Guiding Generative Recommenders with Structured Human Priors via Multi-Head Decoding

Yunkai Zhang, Qiang Zhang, Feng Lin +7

Optimizing recommender systems for objectives beyond accuracy, such as diversity, novelty, and personalization, is crucial for long-term user satisfaction. To this end, industrial…

cs.CV2024

CompCap: Improving Multimodal Large Language Models with Composite Captions

Xiaohui Chen, Satya Narayan Shukla, Mahmoud Azab +8

How well can Multimodal Large Language Models (MLLMs) understand composite images? Composite images (CIs) are synthetic visuals created by merging multiple visual elements, such as…

eess.IV2020

Measure Anatomical Thickness from Cardiac MRI with Deep Neural Networks

Qiaoying Huang, Eric Z. Chen, Hanchao Yu +4

Accurate estimation of shape thickness from medical images is crucial in clinical applications. For example, the thickness of myocardium is one of the key to cardiac disease diagno…

cs.AI2026

SVR-R1: Bootstrapping Multi-modal Reasoning with Self-verification in Reinforcement Learning

Mingyuan Wu, Jingcheng Yang, Shengyi Qian +11

The paper introduces SVR-R1, a reinforcement learning framework that lets a multimodal model generate an answer and then self‑verify it with a binary verdict, allowing a second‑cha…

#multimodal reasoning#reinforcement learning#self-verification#vision-language models
cs.CV2023

SVT: Supertoken Video Transformer for Efficient Video Understanding

Chenbin Pan, Rui Hou, Hanchao Yu +3

Whether by processing videos with fixed resolution from start to end or incorporating pooling and down-scaling strategies, existing video transformers process the whole video conte…

cs.IR2026

RecoWorld: Building Simulated Environments for Agentic Recommender Systems

Fei Liu, Xinyu Lin, Hanchao Yu +12

We present RecoWorld, a blueprint for building simulated environments tailored to agentic recommender systems. Such environments give agents a proper training space where they can…

cs.CL2026

Detecting AI-Generated Content on Social Media with Multi-modal Language Models

Chenyang Yang, Shen Yan, Yibo Yang +13

Generative AI has enabled the creation of photorealistic images and videos that are increasingly disseminated on social media, often used for spam, misinformation, manipulation, an…

cs.AI2026

TTE-Flash: Accelerating Reasoning-based Multimodal Representations via Think-Then-Embed Tokens

Jianpeng Cheng, Xian Wu, Jiangfan Zhang +10

Recent research has demonstrated that Universal Multimodal Embedding (UME) benefits significantly from Chain-of-Thought (CoT) reasoning. In this paradigm, a generative model produc…

cs.IR2025

Towards An Efficient LLM Training Paradigm for CTR Prediction

Allen Lin, Renqin Cai, Yun He +5

Large Language Models (LLMs) have demonstrated tremendous potential as the next-generation ranking-based recommendation system. Many recent works have shown that LLMs can significa…

cs.LG2026

VTool-R1: VLMs Learn to Think with Images via Reinforcement Learning on Multimodal Tool Use

Mingyuan Wu, Jingcheng Yang, Jize Jiang +6

Reinforcement Learning Finetuning (RFT) has significantly advanced the reasoning capabilities of large language models (LLMs) by enabling long chains of thought, self-correction, a…

cs.CL2025

PersonaMem-v2: Towards Personalized Intelligence via Learning Implicit User Personas and Agentic Memory

Bowen Jiang, Yuan Yuan, Maohao Shen +13

Personalization is one of the next milestones in advancing AI capability and alignment. We introduce PersonaMem-v2, the state-of-the-art dataset for LLM personalization that simula…

cs.CV2025

CAFe: Unifying Representation and Generation with Contrastive-Autoregressive Finetuning

Hao Yu, Zhuokai Zhao, Shen Yan +7

The rapid advancement of large vision-language models (LVLMs) has driven significant progress in multimodal tasks, enabling models to interpret, reason, and generate outputs across…

cs.IR2025

Optimizing Recall or Relevance? A Multi-Task Multi-Head Approach for Item-to-Item Retrieval in Recommendation

Jiang Zhang, Sumit Kumar, Wei Chang +7

The task of item-to-item (I2I) retrieval is to identify a set of relevant and highly engaging items based on a given trigger item. It is a crucial component in modern recommendatio…

cs.CV2025

Inference Compute-Optimal Video Vision Language Models

Peiqi Wang, ShengYun Peng, Xuewen Zhang +5

This work investigates the optimal allocation of inference compute across three key scaling factors in video vision language models: language model size, frame count, and the numbe…

cs.CV2020

FOAL: Fast Online Adaptive Learning for Cardiac Motion Estimation

Hanchao Yu, Shanhui Sun, Haichao Yu +4

Motion estimation of cardiac MRI videos is crucial for the evaluation of human heart anatomy and function. Recent researches show promising results with deep learning-based methods…

cs.AI2026

Spreadsheet-RL: Advancing Large Language Model Agents on Realistic Spreadsheet Tasks via Reinforcement Learning

Banghao Chi, Yining Xie, Mingyuan Wu +9

Spreadsheet systems (e.g., Microsoft Excel, Google Sheets) play a central role in modern data-centric workflows. As AI agents grow increasingly capable of automating complex tasks,…

cs.IR2025

Reason to Contrast: A Cascaded Multimodal Retrieval Framework

Xuanming Cui, Hong-You Chen, Hao Yu +10

Traditional multimodal retrieval systems rely primarily on bi-encoder architectures, where performance is closely tied to embedding dimensionality. Recent work, Think-Then-Embed (T…

eess.IV2020

Anatomy-Aware Cardiac Motion Estimation

Pingjun Chen, Xiao Chen, Eric Z. Chen +3

Cardiac motion estimation is critical to the assessment of cardiac function. Myocardium feature tracking (FT) can directly estimate cardiac motion from cine MRI, which requires no…

cs.LG2026

Aha Moment Revisited: Are VLMs Truly Capable of Self Verification in Inference-time Scaling?

Mingyuan Wu, Meitang Li, Jingcheng Yang +6

Inference time techniques such as decoding time scaling and self refinement have been shown to substantially improve mathematical reasoning in large language models (LLMs), largely…