5 papers
Web-CogReasoner: Towards Multimodal Knowledge-Induced Cognitive Reasoning for Web Agents
Yuhan Guo, Cong Guo, Aiwen Sun +12
Multimodal large-scale models have significantly advanced the development of web agents, enabling perception and interaction with digital environments akin to human cognition. In t…
PMDformer: Patch-Mean Decoupling Information Transformer for Long-term Forecasting
Ao Hu, Liangjian Wen, Jiang Duan +7
Long-term time series forecasting (LTSF) plays a crucial role in fields such as energy management, finance, and traffic prediction. Transformer-based models have adopted patch-base…
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models
Faqiang Qian, Kang An, Weikun Zhang +6
Post-training alignment of large language models often combines supervised fine-tuning (SFT) on expert demonstrations with reinforcement learning (RL) from preference or verifiable…
InfMasking: Unleashing Synergistic Information by Contrastive Multimodal Interactions
Liangjian Wen, Qun Dai, Jianzhuang Liu +7
In multimodal representation learning, synergistic interactions between modalities not only provide complementary information but also create unique outcomes through specific inter…
Nexus: An Omni-Perceptive And -Interactive Model for Language, Audio, And Vision
Che Liu, Yingji Zhang, Dong Zhang +13
This work proposes an industry-level omni-modal large language model (LLM) pipeline that integrates auditory, visual, and linguistic modalities to overcome challenges such as limit…