5 papers
Leveraging Vision-Language Large Models for Interpretable Video Action Recognition with Semantic Tokenization
Jingwei Peng, Zhixuan Qiu, Boyu Jin +1
Human action recognition often struggles with deep semantic understanding, complex contextual information, and fine-grained distinction, limitations that traditional methods freque…
MOCHA: Discovering Multi-Order Dynamic Causality in Temporal Point Processes
Yunyang Cao, Juekai Lin, Wenhao Li +1
Discovering complex causal dependencies in temporal point processes (TPPs) is critical for modeling real-world event sequences. Existing methods typically rely on static or first-o…
TextAtari: 100K Frames Game Playing with Language Agents
Wenhao Li, Wenwu Li, Chuyun Shen +8
We present TextAtari, a benchmark for evaluating language agents on very long-horizon decision-making tasks spanning up to 100,000 steps. By translating the visual state representa…
Reinforced Reasoning for Embodied Planning
Di Wu, Jiaxin Fan, Junzhe Zang +4
Embodied planning requires agents to make coherent multi-step decisions based on dynamic visual observations and natural language goals. While recent vision-language models (VLMs)…
Interpretable Hybrid-Rule Temporal Point Processes
Yunyang Cao, Juekai Lin, Hongye Wang +2
Temporal Point Processes (TPPs) are widely used for modeling event sequences in various medical domains, such as disease onset prediction, progression analysis, and clinical decisi…