32 papers
Failure-Informed Image Self-Augmentation for Multimodal Large Language Model Self-Improvement
Chunyang Jiang, Pingping Zhang, Yuzhi Zhao +9
Multimodal large language models (MLLMs) have achieved remarkable performance across vision-language tasks, but their progress depends heavily on large-scale, high-quality multimod…
Universal Image Restoration via Internalized Chain-of-Thought Reasoning
Yu Guo, Zhengru Fang, Shengfeng He +4
Image restoration seeks to recover high-quality images from degraded inputs but becomes highly ill-posed under complex, mixed degradations. While unified all-in-one models are comm…
Multi-Agent Embodied Autonomous Driving (MAEAD): From V2X Information Exchange to Shared World Models
Senkang Hu, Zhengru Fang, Yihang Tao +4
Autonomous driving is shifting from isolated vehicle intelligence toward multi-agent embodied systems that share perception, infer intent, and coordinate action under uncertainty.…
Optimizing Agentic Reasoning with Retrieval via Synthetic Semantic Information Gain Reward
Senkang Hu, Yong Dai, Yuzhi Zhao +5
Agentic reasoning enables large reasoning models (LRMs) to dynamically acquire external knowledge, but yet optimizing the retrieval process remains challenging due to the lack of d…
Unified Context Evolution for LLM Agents
Zixuan Zhu, Yitong Hu, Yong Dai +4
LLM-based agents can solve multi-step interactive tasks by combining reasoning with environment feedback, yet each episode starts from the same fixed context and any useful strateg…
V2VCrafter: Consistent Street-View Image Generation Across Vehicles
Yihang Tao, Yu Guo, Senkang Hu +4
Connected and autonomous driving (CAD) systems leverage vehicle-to-vehicle (V2V) communication for multi-agent collaborative perception, yet remain constrained by scarce annotated…