2 papers
cs.AI2026
Athena-Brain Technical Report: An Efficient Robot Brain for General Intelligence and Embodied Interaction
Jialian Li, Junhong Liu, Yuchen Cao +6
Large language models (LLMs) have demonstrated remarkable capabilities in language understanding, reasoning, and world knowledge. As embodied agents become increasingly capable, th…
cs.CV2026
Video-Text Temporal Localization via Multi-Scale Convolution and Dynamic Routing
Gengtian Shi, Jinze Yu, Chenhao Wu +5
Video-text temporal localization requires precise alignment between natural language queries and corresponding video segments, a fundamental challenge in multimodal understanding.…