3 papers
cs.AI2026
Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models
Zhi Zeng, Cheng Zhang, Zesheng Yang +9
Understanding dynamic sound sources requires jointly determining what produces a sound, where the source is located, and how it moves over time. Yet existing audio-language models…
cs.AI2025
ActionFlow: A Pipelined Action Acceleration for Vision Language Models on Edge
Yuntao Dai, Hang Gu, Teng Wang +6
Vision-Language-Action (VLA) models have emerged as a unified paradigm for robotic perception and control, enabling emergent generalization and long-horizon task execution. However…
cs.CY2025
A Review of Generative AI in Computer Science Education: Challenges and Opportunities in Accuracy, Authenticity, and Assessment
Iman Reihanian, Yunfei Hou, Yu Chen +1
This paper surveys the use of Generative AI tools, such as ChatGPT and Claude, in computer science education, focusing on key aspects of accuracy, authenticity, and assessment. Thr…