9 papers
Learning Spatiotemporal Sensitivity in Video LLMs via Counterfactual Reinforcement Learning
Dazhao Du, Jian Liu, Jialong Qin +7
Video large language models (Video LLMs) achieve strong benchmark accuracy, yet often answer video questions through shortcuts such as single-frame cues and language priors rather…
MLLMs Know When Before Speaking: Revealing and Recovering Temporal Grounding via Attention Cues
Dazhao Du, Liao Duan, Jian Liu +5
Video temporal grounding (VTG), which localizes the start and end times of a queried event in an untrimmed video, is a key test of whether multimodal large language models (MLLMs)…
Robo-Cortex: A Self-Evolving Embodied Agent via Dual-Grain Cognitive Memory and Autonomous Knowledge Induction
Nga Teng Chan, Yi Zhang, Yechi Liu +9
The ability to navigate and interact with complex environments is central to real-world embodied agents, yet navigation in unseen environments remains challenging due to "experient…
BLM: A Boundless Large Model for Cross-Space, Cross-Task, and Cross-Embodiment Learning
Wentao Tan, Bowen Wang, Heng Zhi +15
Multimodal large language models (MLLMs) have advanced vision-language reasoning and are increasingly deployed in embodied agents. However, significant limitations remain: MLLMs ge…
Every Step Evolves: Scaling Reinforcement Learning for Trillion-Scale Thinking Model
Ling Team, Anqi Shen, Baihui Li +101
We present Ring-1T, the first open-source, state-of-the-art thinking model with a trillion-scale parameter. It features 1 trillion total parameters and activates approximately 50 b…
Multi-branch of Attention Yields Accurate Results for Tabular Data
Xuechen Li, Yupeng Li, Jian Liu +2
Tabular data inherently exhibits significant feature heterogeneity, but existing transformer-based methods lack specialized mechanisms to handle this property. To bridge the gap, w…