4 papers
Web-CogReasoner: Towards Multimodal Knowledge-Induced Cognitive Reasoning for Web Agents
Yuhan Guo, Cong Guo, Aiwen Sun +12
Multimodal large-scale models have significantly advanced the development of web agents, enabling perception and interaction with digital environments akin to human cognition. In t…
PMDformer: Patch-Mean Decoupling Information Transformer for Long-term Forecasting
Ao Hu, Liangjian Wen, Jiang Duan +7
Long-term time series forecasting (LTSF) plays a crucial role in fields such as energy management, finance, and traffic prediction. Transformer-based models have adopted patch-base…
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models
Faqiang Qian, Kang An, Weikun Zhang +6
Post-training alignment of large language models often combines supervised fine-tuning (SFT) on expert demonstrations with reinforcement learning (RL) from preference or verifiable…
GazeCLIP: Enhancing Gaze Estimation Through Text-Guided Multimodal Learning
Jun Wang, Hao Ruan, Liangjian Wen +2
Visual gaze estimation, with its wide-ranging application scenarios, has garnered increasing attention within the research community. Although existing approaches infer gaze solely…