7 papers
FeedbackTrack: Visual-Cortex-Inspired Cross-Frame Feedback for Transformer Tracking
Yueyang Cang, Xiaoteng Zhang, Zhiyuan Ning +2
Visual object tracking requires effective temporal integration, yet most Transformer trackers still rely on predominantly feed-forward feature extraction. Existing temporal mechani…
SAW: Stage-Aware Dynamic Weighting for Multi-Objective Reinforcement Learning in Large Language Models
Yuchen He, Baolong Bi, Shenghua Liu +7
Although multi-objective reinforcement learning (MORL) is central to aligning large language models with complex human preferences, the prevailing practice of static weighted summa…
PRIMED: Adaptive Modality Suppression for Referring Audio-Visual Segmentation via Biased Competition
Yuchen He, Jing Zhang
Referring Audio-Visual Segmentation (Ref-AVS) seeks to localize and segment target objects in video frames based on visual, auditory, and textual referring cues. The task is challe…
Structure-Guided Diffusion Model for EEG-Based Visual Cognition Reconstruction
Yongxiang Lian, Yueyang Cang, Pingge Hu +2
Objective: Decoding visual information from electroencephalography (EEG) is an important problem in neuroscience and brain-computer interface (BCI) research. Existing methods are l…
Graph-GRPO: Stabilizing Multi-Agent Topology Learning via Group Relative Policy Optimization
Yueyang Cang, Xiaoteng Zhang, Erlu Zhao +7
Optimizing communication topology is fundamental to the efficiency and effectiveness of Large Language Model (LLM)-based Multi-Agent Systems (MAS). While recent approaches utilize…
PromptCD: Test-Time Behavior Enhancement via Polarity-Prompt Contrastive Decoding
Baolong Bi, Yuyao Ge, Shenghua Liu +9
Reliable AI systems require large language models (LLMs) to exhibit behaviors aligned with human preferences and values. However, most existing alignment approaches operate at trai…