21 papers · 1 filter
From Seeing to Acting: Smart Glasses as First-Person Intelligence Platforms
Jiangning Zhang, Haojun Chen, Yong Liu
Smart glasses are evolving from capture and display accessories into first-person intelligence platforms that connect human perception, persistent context, and digital or physical…
Deep Academic Survey: Stateful Agentic Closed-Loop Paradigm for Academic Survey Automation
Zhikai Xu, Zhucun Xue, Teng Hu +3
Academic surveys play a central role in organizing rapidly expanding scholarly literature, yet their construction requires extensive paper analysis, coherent knowledge organization…
Quo Vadis, World Modeling?
Yu Yang, Xuemeng Yang, Licheng Wen +17
Continually improving agents require dynamic interaction feedback beyond static supervision, yet direct real-environment interaction is costly, slow, unsafe, and hard to paralleliz…
ChatImage: Navigating Long-Form LLM Answers through Interactive Images
Wencan Jiang, Jiangning Zhang, Yong Liu
Large Language Models (LLMs) can produce detailed answers to complex queries, but these answers are typically presented as dense linear text, which makes fine-grained inspection, n…
SPOT-E: Test-Time Entropy Shaping with Visual Spotlights for Frozen VLMs
Bo Yin, Xiaobin Hu, Chengming Xu +6
Vision-language models (VLMs) often underperform on evidence intensive tasks because decisive visual evidence are small, localized, and easy to overlook, leading to failures in evi…
What Semantics Survive the Connector? Diagnosing VLM-to-DiT Alignment in Video Editing
Hangyu Lin, Chao Wen, Chengming Xu +4
Flow matching based video generative models have been increasingly relying on prepended Vision-Language Models (VLMs) to handle complex, instruction-based video editing. The prevai…