8 papers
AVA-Encoder: Towards Agent-Native Video Representation Learning
Chuyue Li, Jinpeng Yu, Haozhe Wang +8
Video creative agents still lack an effective way to learn from high-quality human films, limiting their ability to produce cinematic-grade videos. A key challenge is the absence o…
DiffImaginE: Imagine to Verify Entity Types with Diffusion
Feng Zhang, Feiyu Han, Rongxin Yang +11
Multimodal named entity recognition (MNER) determines whether each candidate span and entity-type hypothesis is supported by joint textual and visual evidence. Existing imagine-and…
Chain-of-Memory: Lightweight Memory Construction with Dynamic Evolution for LLM Agents
Xiucheng Xu, Bingbing Xu, Xueyun Tian +4
External memory systems are pivotal for enabling Large Language Model (LLM) agents to maintain persistent knowledge and perform long-horizon decision-making. Existing paradigms typ…
Towards Robust Process Reward Modeling via Noise-aware Learning
Bin Xie, Bingbing Xu, Xueyun Tian +2
Process Reward Models (PRMs) have achieved strong results in complex reasoning, but are bottlenecked by costly process-level supervision. A widely used alternative, Monte Carlo Est…
ROMA: Real-time Omni-Multimodal Assistant with Interactive Streaming Understanding
Xueyun Tian, Wei Li, Bingbing Xu +3
Recent Omni-multimodal Large Language Models show promise in unified audio, vision, and text modeling. However, streaming audio-video understanding remains challenging, as existing…
Beyond Entangled Planning: Task-Decoupled Planning for Long-Horizon Agents
Yunfan Li, Bingbing Xu, Xueyun Tian +2
Recent advances in large language models (LLMs) have enabled agents to autonomously execute complex, long-horizon tasks, yet planning remains a primary bottleneck for reliable task…