Publications (30)
Efficient Sequential Recommendation for Long Term User Interest Via Personalization
Qiang Zhang, Hanchao Yu, Ivan Ji +14
Recent years have witnessed success of sequential modeling, generative recommender, and large language model for recommendation. Though the scaling law has been validated for seque…
BRIDLE: Generalized Self-supervised Learning with Quantization
Hoang M. Nguyen, Satya N. Shukla, Qiang Zhang +5
Self-supervised learning has been a powerful approach for learning meaningful representations from unlabeled data across various domains, reducing the reliance on large labeled dat…
Learning Critically: Selective Self Distillation in Federated Learning on Non-IID Data
Yuting He, Yiqiang Chen, XiaoDong Yang +3
Federated learning (FL) enables multiple clients to collaboratively train a global model while keeping local data decentralized. Data heterogeneity (non-IID) across clients has imp…
GISTBench: Evaluating LLM User Understanding via Evidence-Based Interest Verification
Iordanis Fostiropoulos, Muhammad Rafay Azhar, Abdalaziz Sawwan +8
We introduce GISTBench, a benchmark for evaluating Large Language Models' (LLMs) ability to understand users from their interaction histories in recommendation systems. Unlike trad…
Study Group Learning: Improving Retinal Vessel Segmentation Trained with Noisy Labels
Yuqian Zhou, Hanchao Yu, Humphrey Shi
Retinal vessel segmentation from retinal images is an essential task for developing the computer-aided diagnosis system for retinal diseases. Efforts have been made on high-perform…
Uniform Masking Prevails in Vision-Language Pretraining
Siddharth Verma, Yuchen Lu, Rui Hou +4
Masked Language Modeling (MLM) has proven to be an essential component of Vision-Language (VL) pretraining. To implement MLM, the researcher must make two design choices: the maski…
Verifiable Reasoning for LLM-based Generative Recommendation
Xinyu Lin, Hanqing Zeng, Hanchao Yu +8
Reasoning in Large Language Models (LLMs) has recently shown strong potential in enhancing generative recommendation through deep understanding of complex user preference. Existing…
MMViT: Multiscale Multiview Vision Transformers
Yuchen Liu, Natasha Ong, Kaiyan Peng +8
We present Multiscale Multiview Vision Transformers (MMViT), which introduces multiscale feature maps and multiview encodings to transformer models. Our model encodes different vie…
Motion Pyramid Networks for Accurate and Efficient Cardiac Motion Estimation
Hanchao Yu, Xiao Chen, Humphrey Shi +3
Cardiac motion estimation plays a key role in MRI cardiac feature tracking and function assessment such as myocardium strain. In this paper, we propose Motion Pyramid Networks, a n…
RoAST: Robustifying Language Models via Adversarial Perturbation with Selective Training
Jaehyung Kim, Yuning Mao, Rui Hou +7
Fine-tuning pre-trained language models (LMs) has become the de facto standard in many NLP tasks. Nevertheless, fine-tuned LMs are still prone to robustness issues, such as adversa…
RESTRAIN: From Spurious Votes to Signals -- Self-Driven RL with Self-Penalization
Zhaoning Yu, Will Su, Leitian Tao +9
Reinforcement learning with human-annotated data has boosted chain-of-thought reasoning in large reasoning models, but these gains come at high costs in labeled data while falterin…
Don't Waste It: Guiding Generative Recommenders with Structured Human Priors via Multi-Head Decoding
Yunkai Zhang, Qiang Zhang, Feng Lin +7
Optimizing recommender systems for objectives beyond accuracy, such as diversity, novelty, and personalization, is crucial for long-term user satisfaction. To this end, industrial…
CompCap: Improving Multimodal Large Language Models with Composite Captions
Xiaohui Chen, Satya Narayan Shukla, Mahmoud Azab +8
How well can Multimodal Large Language Models (MLLMs) understand composite images? Composite images (CIs) are synthetic visuals created by merging multiple visual elements, such as…
Measure Anatomical Thickness from Cardiac MRI with Deep Neural Networks
Qiaoying Huang, Eric Z. Chen, Hanchao Yu +4
Accurate estimation of shape thickness from medical images is crucial in clinical applications. For example, the thickness of myocardium is one of the key to cardiac disease diagno…
SVR-R1: Bootstrapping Multi-modal Reasoning with Self-verification in Reinforcement Learning
Mingyuan Wu, Jingcheng Yang, Shengyi Qian +11
The paper introduces SVR-R1, a reinforcement learning framework that lets a multimodal model generate an answer and then self‑verify it with a binary verdict, allowing a second‑cha…
SVT: Supertoken Video Transformer for Efficient Video Understanding
Chenbin Pan, Rui Hou, Hanchao Yu +3
Whether by processing videos with fixed resolution from start to end or incorporating pooling and down-scaling strategies, existing video transformers process the whole video conte…
RecoWorld: Building Simulated Environments for Agentic Recommender Systems
Fei Liu, Xinyu Lin, Hanchao Yu +12
We present RecoWorld, a blueprint for building simulated environments tailored to agentic recommender systems. Such environments give agents a proper training space where they can…
Detecting AI-Generated Content on Social Media with Multi-modal Language Models
Chenyang Yang, Shen Yan, Yibo Yang +13
Generative AI has enabled the creation of photorealistic images and videos that are increasingly disseminated on social media, often used for spam, misinformation, manipulation, an…
TTE-Flash: Accelerating Reasoning-based Multimodal Representations via Think-Then-Embed Tokens
Jianpeng Cheng, Xian Wu, Jiangfan Zhang +10
Recent research has demonstrated that Universal Multimodal Embedding (UME) benefits significantly from Chain-of-Thought (CoT) reasoning. In this paradigm, a generative model produc…
Towards An Efficient LLM Training Paradigm for CTR Prediction
Allen Lin, Renqin Cai, Yun He +5
Large Language Models (LLMs) have demonstrated tremendous potential as the next-generation ranking-based recommendation system. Many recent works have shown that LLMs can significa…
VTool-R1: VLMs Learn to Think with Images via Reinforcement Learning on Multimodal Tool Use
Mingyuan Wu, Jingcheng Yang, Jize Jiang +6
Reinforcement Learning Finetuning (RFT) has significantly advanced the reasoning capabilities of large language models (LLMs) by enabling long chains of thought, self-correction, a…
PersonaMem-v2: Towards Personalized Intelligence via Learning Implicit User Personas and Agentic Memory
Bowen Jiang, Yuan Yuan, Maohao Shen +13
Personalization is one of the next milestones in advancing AI capability and alignment. We introduce PersonaMem-v2, the state-of-the-art dataset for LLM personalization that simula…
CAFe: Unifying Representation and Generation with Contrastive-Autoregressive Finetuning
Hao Yu, Zhuokai Zhao, Shen Yan +7
The rapid advancement of large vision-language models (LVLMs) has driven significant progress in multimodal tasks, enabling models to interpret, reason, and generate outputs across…
Optimizing Recall or Relevance? A Multi-Task Multi-Head Approach for Item-to-Item Retrieval in Recommendation
Jiang Zhang, Sumit Kumar, Wei Chang +7
The task of item-to-item (I2I) retrieval is to identify a set of relevant and highly engaging items based on a given trigger item. It is a crucial component in modern recommendatio…
Inference Compute-Optimal Video Vision Language Models
Peiqi Wang, ShengYun Peng, Xuewen Zhang +5
This work investigates the optimal allocation of inference compute across three key scaling factors in video vision language models: language model size, frame count, and the numbe…
FOAL: Fast Online Adaptive Learning for Cardiac Motion Estimation
Hanchao Yu, Shanhui Sun, Haichao Yu +4
Motion estimation of cardiac MRI videos is crucial for the evaluation of human heart anatomy and function. Recent researches show promising results with deep learning-based methods…
Spreadsheet-RL: Advancing Large Language Model Agents on Realistic Spreadsheet Tasks via Reinforcement Learning
Banghao Chi, Yining Xie, Mingyuan Wu +9
Spreadsheet systems (e.g., Microsoft Excel, Google Sheets) play a central role in modern data-centric workflows. As AI agents grow increasingly capable of automating complex tasks,…
Reason to Contrast: A Cascaded Multimodal Retrieval Framework
Xuanming Cui, Hong-You Chen, Hao Yu +10
Traditional multimodal retrieval systems rely primarily on bi-encoder architectures, where performance is closely tied to embedding dimensionality. Recent work, Think-Then-Embed (T…
Anatomy-Aware Cardiac Motion Estimation
Pingjun Chen, Xiao Chen, Eric Z. Chen +3
Cardiac motion estimation is critical to the assessment of cardiac function. Myocardium feature tracking (FT) can directly estimate cardiac motion from cine MRI, which requires no…
Aha Moment Revisited: Are VLMs Truly Capable of Self Verification in Inference-time Scaling?
Mingyuan Wu, Meitang Li, Jingcheng Yang +6
Inference time techniques such as decoding time scaling and self refinement have been shown to substantially improve mathematical reasoning in large language models (LLMs), largely…