activity
20242026
collaborators

11 papers

cs.CV2026

DesignAgent3D: Interactive 3D Scene Editing via Designer-like Multimodal Reasoning

Xiujin Liu, Tianyu Yang, Yilun Zhao +1

Text guided 3D scene editing provides an intuitive interface for modifying reconstructed environments, but remains difficult because natural language design requests are often sema…

cs.AI2026

Reason Before You Retrieve: Agentic Planning for Multi-modal RAG

Tianyu Yang, Shir Simon, Zhenzhen Li +2

Multimodal retrieval-augmented generation (mRAG) aims to answer image-text queries with external knowledge, but most existing systems still retrieve directly from raw multimodal in…

cs.LG2026

Getting Better at Working With You: Compiling User Corrections into Runtime Enforcement for Coding Agents

Yujun Zhou, Kehan Guo, Haomin Zhuang +8

Interactive LLM agents are becoming part of daily work, but they do not reliably become easier to work with over time: a correction remembered in one session may still be violated…

cs.AI2026

A Survey of Multimodal Mathematical Reasoning: From Perception, Alignment to Reasoning

Tianyu Yang, Sihong Wu, Yilun Zhao +6

Multimodal Mathematical Reasoning (MMR) has recently attracted increasing attention for its capability to solve mathematical problems involving both textual and visual modalities.…

cs.CY2025

The Role of Computing Resources in Publishing Foundation Model Research

Yuexing Hao, Yue Huang, Haoran Zhang +8

Cutting-edge research in Artificial Intelligence (AI) requires considerable resources, including Graphics Processing Units (GPUs), data, and human resources. In this paper, we eval…

cs.LG2025

Evolving Language Models without Labels: Majority Drives Selection, Novelty Promotes Variation

Yujun Zhou, Zhenwen Liang, Haolin Liu +7

Large language models (LLMs) are increasingly trained with reinforcement learning from verifiable rewards (RLVR), yet real-world deployment demands models that can self-improve wit…