Publications (30)
TWIN V2: Scaling Ultra-Long User Behavior Sequence Modeling for Enhanced CTR Prediction at Kuaishou
Zihua Si, Lin Guan, ZhongXiang Sun +12
The significance of modeling long-term user interests for CTR prediction tasks in large-scale recommendation systems is progressively gaining attention among researchers and practi…
Robust Planning with LLM-Modulo Framework: Case Study in Travel Planning
Atharva Gundawar, Mudit Verma, Lin Guan +3
As the applicability of Large Language Models (LLMs) extends beyond traditional text processing tasks, there is a burgeoning interest in their potential to excel in planning and re…
Relative Behavioral Attributes: Filling the Gap between Symbolic Goal Specification and Reward Learning from Human Preferences
Lin Guan, Karthik Valmeekam, Subbarao Kambhampati
Generating complex behaviors that satisfy the preferences of non-expert users is a crucial requirement for AI agents. Interactive reward learning from trajectory comparisons (a.k.a…
LLMs Can't Plan, But Can Help Planning in LLM-Modulo Frameworks
Subbarao Kambhampati, Karthik Valmeekam, Lin Guan +5
There is considerable confusion about the role of Large Language Models (LLMs) in planning and reasoning tasks. On one side are over-optimistic claims that LLMs can indeed do these…
TDMP-Reliable Target Driven and Mobility Prediction based Routing Protocol in Complex VANET
Mao Ye, Lin Guan, Mohammed Quddus
Vehicle-to-everything (V2X) communication in the vehicular ad hoc network (VANET), an infrastructure-free mechanism, has emerged as a crucial component in the advanced Intelligent…
Leveraging Approximate Symbolic Models for Reinforcement Learning via Skill Diversity
Lin Guan, Sarath Sreedharan, Subbarao Kambhampati
Creating reinforcement learning (RL) agents that are capable of accepting and leveraging task-specific knowledge from humans has been long identified as a possible strategy for dev…
Enhanced Exploration in Neural Feature Selection for Deep Click-Through Rate Prediction Models via Ensemble of Gating Layers
Lin Guan, Xia Xiao, Ming Chen +1
Feature selection has been an essential step in developing industry-scale deep Click-Through Rate (CTR) prediction systems. The goal of neural feature selection (NFS) is to choose…
Extracting Heuristics from Large Language Models for Reward Shaping in Reinforcement Learning
Siddhant Bhambri, Amrita Bhattacharjee, Durgesh Kalwar +3
Reinforcement Learning (RL) suffers from sample inefficiency in sparse reward domains, and the problem is further pronounced in case of stochastic transitions. To improve the sampl…
Task Success is not Enough: Investigating the Use of Video-Language Models as Behavior Critics for Catching Undesirable Agent Behaviors
Lin Guan, Yifan Zhou, Denis Liu +3
Large-scale generative models are shown to be useful for sampling meaningful candidate solutions, yet they often overlook task constraints and user preferences. Their full power is…
QoS-aware Link Scheduling Strategy for Data Transmission in SDVN
Yong Zhang, Mao Ye, Lin Guan
The vehicular ad-hoc network (VANET) based on dedicated short-range communication (DSRC) is a distributed communication system, in which all the nodes share the wireless channel wi…
Paying Less Generalization Tax: A Cross-Domain Generalization Study of RL Training for LLM Agents
Zhihan Liu, Lin Guan, Yixin Nie +6
Generalist LLM agents are often post-trained on a narrow set of environments but deployed across far broader, unseen domains. In this work, we investigate the challenge of agentic…
Bi-Josephson Effect in a Driven-Dissipative Supersolid
Jieli Qin, Shijie Li, Yijia Tu +4
The Josephson effect is a macroscopic quantum tunneling phenomenon in a system with superfluid property, when it is split into two parts by a barrier. Here, we examine the Josephso…
A Product Line Systems Engineering Process for Variability Identification and Reduction
Mole Li, Alan Grigg, Charles Dickerson +2
Software Product Line Engineering has attracted attention in the last two decades due to its promising capabilities to reduce costs and time to market through reuse of requirements…
Atari-HEAD: Atari Human Eye-Tracking and Demonstration Dataset
Ruohan Zhang, Calen Walshe, Zhuode Liu +6
Large-scale public datasets have been shown to benefit research in multiple areas of modern artificial intelligence. For decision-making research that requires human data, high-qua…
Make It Long, Keep It Fast: End-to-End 10K Long User Behavior Sequence Modeling for Billion-Scale Douyin Recommendation
Lin Guan, Jia-Qi Yang, Zhishan Zhao +12
Short-video recommenders such as Douyin must exploit extremely long user behavior histories without breaking latency or cost budgets. We present an end-to-end industrial recommende…
Leveraging Human Guidance for Deep Reinforcement Learning Tasks
Ruohan Zhang, Faraz Torabi, Lin Guan +2
Reinforcement learning agents can learn to solve sequential decision tasks by interacting with the environment. Human knowledge of how to solve these tasks can be incorporated usin…
Widening the Pipeline in Human-Guided Reinforcement Learning with Explanation and Context-Aware Data Augmentation
Lin Guan, Mudit Verma, Sihang Guo +2
Human explanation (e.g., in terms of feature importance) has been recently used to extend the communication channel between human and agent in interactive machine learning. Under t…
CharacterFlywheel: Scaling Iterative Improvement of Engaging and Steerable LLMs in Production
Yixin Nie, Lin Guan, Zhongyao Ma +19
This report presents CharacterFlywheel, an iterative flywheel process for improving large language models (LLMs) in production social chat applications across Instagram, WhatsApp,…
Pyramid Mixer: Multi-dimensional Multi-period Interest Modeling for Sequential Recommendation
Zhen Gong, Zhifang Fan, Hui Lu +7
Sequential recommendation, a critical task in recommendation systems, predicts the next user action based on the understanding of the user's historical behaviors. Conventional stud…
Symbols as a Lingua Franca for Bridging Human-AI Chasm for Explainable and Advisable AI Systems
Subbarao Kambhampati, Sarath Sreedharan, Mudit Verma +2
Despite the surprising power of many modern AI systems that often learn their own representations, there is significant discontent about their inscrutability and the attendant prob…
Leveraging Pre-trained Large Language Models to Construct and Utilize World Models for Model-based Task Planning
Lin Guan, Karthik Valmeekam, Sarath Sreedharan +1
There is a growing interest in applying pre-trained large language models (LLMs) to planning problems. However, methods that use LLMs directly as planners are currently impractical…
Towards customizable reinforcement learning agents: Enabling preference specification through online vocabulary expansion
Utkarsh Soni, Nupur Thakur, Sarath Sreedharan +4
There is a growing interest in developing automated agents that can work alongside humans. In addition to completing the assigned task, such an agent will undoubtedly be expected t…
Contrastively Learning Visual Attention as Affordance Cues from Demonstrations for Robotic Grasping
Yantian Zha, Siddhant Bhambri, Lin Guan
Conventional works that learn grasping affordance from demonstrations need to explicitly predict grasping configurations, such as gripper approaching angles or grasping preshapes.…
Learning from Ambiguous Demonstrations with Self-Explanation Guided Reinforcement Learning
Yantian Zha, Lin Guan, Subbarao Kambhampati
Our work aims at efficiently leveraging ambiguous demonstrations for the training of a reinforcement learning (RL) agent. An ambiguous demonstration can usually be interpreted in m…
TWIN: TWo-stage Interest Network for Lifelong User Behavior Modeling in CTR Prediction at Kuaishou
Jianxin Chang, Chenbin Zhang, Zhiyi Fu +8
Life-long user behavior modeling, i.e., extracting a user's hidden interests from rich historical behaviors in months or even years, plays a central role in modern CTR prediction s…
HorizonBench: Long-Horizon Personalization with Evolving Preferences
Shuyue Stella Li, Bhargavi Paranjape, Kerem Oktar +9
User preferences evolve across months of interaction, and tracking them requires inferring when a stated preference has been changed by a subsequent life event. We define this prob…
Data-Driven Online Optimization for Enhancing Power System Oscillation Damping
Zhihao Chen, Hanchen Xu, Junbo Zhang +1
This paper reports an initial work on power system oscillation damping improvement using a data-driven online optimization method. An online oscillation damping optimization mod-el…
Overlap-Minimization Scheduling Strategy for Data Transmission in VANET
Yong Zhang, Mao Ye, Lin Guan
The vehicular ad-hoc network (VANET) based on dedicated short-range communication (DSRC) is a distributed communication system, in which all the nodes share the wireless channel wi…
MuSCLe: A Multi-Strategy Contrastive Learning Framework for Weakly Supervised Semantic Segmentation
Kunhao Yuan, Gerald Schaefer, Yu-Kun Lai +4
Weakly supervised semantic segmentation (WSSS) has gained significant popularity since it relies only on weak labels such as image level annotations rather than pixel level annotat…
Scale-Adaptive Power Flow Analysis with Local Topology Slicing and Multi-Task Graph Learning
Yongzhe Li, Lin Guan, Zihan Cai +3
Developing deep learning models with strong adaptability to topological variations is of great practical significance for power flow analysis. To enhance model performance under va…