papers

Publications (30)

cs.IR2024

TWIN V2: Scaling Ultra-Long User Behavior Sequence Modeling for Enhanced CTR Prediction at Kuaishou

Zihua Si, Lin Guan, ZhongXiang Sun +12

The significance of modeling long-term user interests for CTR prediction tasks in large-scale recommendation systems is progressively gaining attention among researchers and practi…

cs.AI2024

Robust Planning with LLM-Modulo Framework: Case Study in Travel Planning

Atharva Gundawar, Mudit Verma, Lin Guan +3

As the applicability of Large Language Models (LLMs) extends beyond traditional text processing tasks, there is a burgeoning interest in their potential to excel in planning and re…

cs.AI2023

Relative Behavioral Attributes: Filling the Gap between Symbolic Goal Specification and Reward Learning from Human Preferences

Lin Guan, Karthik Valmeekam, Subbarao Kambhampati

Generating complex behaviors that satisfy the preferences of non-expert users is a crucial requirement for AI agents. Interactive reward learning from trajectory comparisons (a.k.a…

cs.AI2024

LLMs Can't Plan, But Can Help Planning in LLM-Modulo Frameworks

Subbarao Kambhampati, Karthik Valmeekam, Lin Guan +5

There is considerable confusion about the role of Large Language Models (LLMs) in planning and reasoning tasks. On one side are over-optimistic claims that LLMs can indeed do these…

cs.NI2020

TDMP-Reliable Target Driven and Mobility Prediction based Routing Protocol in Complex VANET

Mao Ye, Lin Guan, Mohammed Quddus

Vehicle-to-everything (V2X) communication in the vehicular ad hoc network (VANET), an infrastructure-free mechanism, has emerged as a crucial component in the advanced Intelligent…

cs.AI2022

Leveraging Approximate Symbolic Models for Reinforcement Learning via Skill Diversity

Lin Guan, Sarath Sreedharan, Subbarao Kambhampati

Creating reinforcement learning (RL) agents that are capable of accepting and leveraging task-specific knowledge from humans has been long identified as a possible strategy for dev…

cs.LG2021

Enhanced Exploration in Neural Feature Selection for Deep Click-Through Rate Prediction Models via Ensemble of Gating Layers

Lin Guan, Xia Xiao, Ming Chen +1

Feature selection has been an essential step in developing industry-scale deep Click-Through Rate (CTR) prediction systems. The goal of neural feature selection (NFS) is to choose…

cs.LG2024

Extracting Heuristics from Large Language Models for Reward Shaping in Reinforcement Learning

Siddhant Bhambri, Amrita Bhattacharjee, Durgesh Kalwar +3

Reinforcement Learning (RL) suffers from sample inefficiency in sparse reward domains, and the problem is further pronounced in case of stochastic transitions. To improve the sampl…

cs.AI2024

Task Success is not Enough: Investigating the Use of Video-Language Models as Behavior Critics for Catching Undesirable Agent Behaviors

Lin Guan, Yifan Zhou, Denis Liu +3

Large-scale generative models are shown to be useful for sampling meaningful candidate solutions, yet they often overlook task constraints and user preferences. Their full power is…

cs.NI2021

QoS-aware Link Scheduling Strategy for Data Transmission in SDVN

Yong Zhang, Mao Ye, Lin Guan

The vehicular ad-hoc network (VANET) based on dedicated short-range communication (DSRC) is a distributed communication system, in which all the nodes share the wireless channel wi…

cs.AI2026

Paying Less Generalization Tax: A Cross-Domain Generalization Study of RL Training for LLM Agents

Zhihan Liu, Lin Guan, Yixin Nie +6

Generalist LLM agents are often post-trained on a narrow set of environments but deployed across far broader, unseen domains. In this work, we investigate the challenge of agentic…

cond-mat.quant-gas2025

Bi-Josephson Effect in a Driven-Dissipative Supersolid

Jieli Qin, Shijie Li, Yijia Tu +4

The Josephson effect is a macroscopic quantum tunneling phenomenon in a system with superfluid property, when it is split into two parts by a barrier. Here, we examine the Josephso…

cs.SE2018

A Product Line Systems Engineering Process for Variability Identification and Reduction

Mole Li, Alan Grigg, Charles Dickerson +2

Software Product Line Engineering has attracted attention in the last two decades due to its promising capabilities to reduce costs and time to market through reuse of requirements…

cs.LG2019

Atari-HEAD: Atari Human Eye-Tracking and Demonstration Dataset

Ruohan Zhang, Calen Walshe, Zhuode Liu +6

Large-scale public datasets have been shown to benefit research in multiple areas of modern artificial intelligence. For decision-making research that requires human data, high-qua…

cs.LG2026

Make It Long, Keep It Fast: End-to-End 10K Long User Behavior Sequence Modeling for Billion-Scale Douyin Recommendation

Lin Guan, Jia-Qi Yang, Zhishan Zhao +12

Short-video recommenders such as Douyin must exploit extremely long user behavior histories without breaking latency or cost budgets. We present an end-to-end industrial recommende…

cs.AI2019

Leveraging Human Guidance for Deep Reinforcement Learning Tasks

Ruohan Zhang, Faraz Torabi, Lin Guan +2

Reinforcement learning agents can learn to solve sequential decision tasks by interacting with the environment. Human knowledge of how to solve these tasks can be incorporated usin…

cs.AI2021

Widening the Pipeline in Human-Guided Reinforcement Learning with Explanation and Context-Aware Data Augmentation

Lin Guan, Mudit Verma, Sihang Guo +2

Human explanation (e.g., in terms of feature importance) has been recently used to extend the communication channel between human and agent in interactive machine learning. Under t…

cs.CL2026

CharacterFlywheel: Scaling Iterative Improvement of Engaging and Steerable LLMs in Production

Yixin Nie, Lin Guan, Zhongyao Ma +19

This report presents CharacterFlywheel, an iterative flywheel process for improving large language models (LLMs) in production social chat applications across Instagram, WhatsApp,…

cs.IR2025

Pyramid Mixer: Multi-dimensional Multi-period Interest Modeling for Sequential Recommendation

Zhen Gong, Zhifang Fan, Hui Lu +7

Sequential recommendation, a critical task in recommendation systems, predicts the next user action based on the understanding of the user's historical behaviors. Conventional stud…

cs.AI2021

Symbols as a Lingua Franca for Bridging Human-AI Chasm for Explainable and Advisable AI Systems

Subbarao Kambhampati, Sarath Sreedharan, Mudit Verma +2

Despite the surprising power of many modern AI systems that often learn their own representations, there is significant discontent about their inscrutability and the attendant prob…

cs.AI2023

Leveraging Pre-trained Large Language Models to Construct and Utilize World Models for Model-based Task Planning

Lin Guan, Karthik Valmeekam, Sarath Sreedharan +1

There is a growing interest in applying pre-trained large language models (LLMs) to planning problems. However, methods that use LLMs directly as planners are currently impractical…

cs.AI2023

Towards customizable reinforcement learning agents: Enabling preference specification through online vocabulary expansion

Utkarsh Soni, Nupur Thakur, Sarath Sreedharan +4

There is a growing interest in developing automated agents that can work alongside humans. In addition to completing the assigned task, such an agent will undoubtedly be expected t…

cs.RO2021

Contrastively Learning Visual Attention as Affordance Cues from Demonstrations for Robotic Grasping

Yantian Zha, Siddhant Bhambri, Lin Guan

Conventional works that learn grasping affordance from demonstrations need to explicitly predict grasping configurations, such as gripper approaching angles or grasping preshapes.…

cs.LG2024

Learning from Ambiguous Demonstrations with Self-Explanation Guided Reinforcement Learning

Yantian Zha, Lin Guan, Subbarao Kambhampati

Our work aims at efficiently leveraging ambiguous demonstrations for the training of a reinforcement learning (RL) agent. An ambiguous demonstration can usually be interpreted in m…

cs.IR2023

TWIN: TWo-stage Interest Network for Lifelong User Behavior Modeling in CTR Prediction at Kuaishou

Jianxin Chang, Chenbin Zhang, Zhiyi Fu +8

Life-long user behavior modeling, i.e., extracting a user's hidden interests from rich historical behaviors in months or even years, plays a central role in modern CTR prediction s…

cs.CL2026

HorizonBench: Long-Horizon Personalization with Evolving Preferences

Shuyue Stella Li, Bhargavi Paranjape, Kerem Oktar +9

User preferences evolve across months of interaction, and tracking them requires inferring when a stated preference has been changed by a subsequent life event. We define this prob…

math.OC2019

Data-Driven Online Optimization for Enhancing Power System Oscillation Damping

Zhihao Chen, Hanchen Xu, Junbo Zhang +1

This paper reports an initial work on power system oscillation damping improvement using a data-driven online optimization method. An online oscillation damping optimization mod-el…

cs.NI2021

Overlap-Minimization Scheduling Strategy for Data Transmission in VANET

Yong Zhang, Mao Ye, Lin Guan

The vehicular ad-hoc network (VANET) based on dedicated short-range communication (DSRC) is a distributed communication system, in which all the nodes share the wireless channel wi…

cs.CV2022

MuSCLe: A Multi-Strategy Contrastive Learning Framework for Weakly Supervised Semantic Segmentation

Kunhao Yuan, Gerald Schaefer, Yu-Kun Lai +4

Weakly supervised semantic segmentation (WSSS) has gained significant popularity since it relies only on weak labels such as image level annotations rather than pixel level annotat…

cs.LG2026

Scale-Adaptive Power Flow Analysis with Local Topology Slicing and Multi-Task Graph Learning

Yongzhe Li, Lin Guan, Zihan Cai +3

Developing deep learning models with strong adaptability to topological variations is of great practical significance for power flow analysis. To enhance model performance under va…