2 papers
cs.CV2025
iMOVE: Instance-Motion-Aware Video Understanding
Jiaze Li, Yaya Shi, Zongyang Ma +7
Enhancing the fine-grained instance spatiotemporal motion perception capabilities of Video Large Language Models is crucial for improving their temporal and general video understan…
cs.CL2024
Kwai-STaR: Transform LLMs into State-Transition Reasoners
Xingyu Lu, Yuhang Hu, Changyi Liu +12
Mathematical reasoning presents a significant challenge to the cognitive capabilities of LLMs. Various methods have been proposed to enhance the mathematical ability of LLMs. Howev…