16 papers
Advantage-Guided Gate: Reshaping Open-Ended Reasoning for Vision-Based Spatial Intelligence
Ling Lin, Yang Bai, Congcong Zhu +6
Multimodal large language models (MLLMs) have demonstrated significant potential in complex spatial scene understanding and reasoning tasks. However, their open-ended reasoning pro…
One Video, One World: Turning Monocular Video into Physical 4D Scenes
Junhao Chen, Boran Zhang, Mingjin Chen +7
We introduce \textbf{OVOW}, the first training-free system that reconstructs \emph{instance-level, simulation-ready} 4D mesh scenes from a single monocular video. Recent 4D reconst…
MeshTok: Efficient Multi-Scale Tokenization for Scalable PDE Transformers
Yanshun Zhao, Xiaoyu Peng, Jiamin Jiang +2
Conventional patchified Transformers operate on uniform spatial partitions, distributing computational effort evenly across the domain irrespective of local features. This inflexib…
From Spark to Fire: Modeling and Mitigating Error Cascades in LLM-Based Multi-Agent Collaboration
Yizhe Xie, Congcong Zhu, Xinyue Zhang +5
Large Language Model-based Multi-Agent Systems (LLM-MAS) are increasingly applied to complex collaborative scenarios. However, their collaborative mechanisms may cause minor inaccu…
When LLMs Team Up: A Coordinated Attack Framework for Automated Cyber Intrusions
Minfeng Qi, Tianqing Zhu, Zijie Xu +3
Automated intrusion-style workflows require LLM agents to reason over partial observations, tool outputs, and executable artifacts under bounded budgets. A single LLM instance ofte…
Learning to Look before Learning to Like: Incorporating Human Visual Cognition into Aesthetic Quality Assessment
Liwen Yu, Chi Liu, Xiaotong Han +3
Automated Aesthetic Quality Assessment (AQA) treats images primarily as static pixel vectors, aligning predictions with human-rating scores largely through semantic perception. How…