activity
20242026
collaborators

10 papers

cs.CV2026

DesignAgent3D: Interactive 3D Scene Editing via Designer-like Multimodal Reasoning

Xiujin Liu, Tianyu Yang, Yilun Zhao +1

Text guided 3D scene editing provides an intuitive interface for modifying reconstructed environments, but remains difficult because natural language design requests are often sema…

cs.AI2026

Reason Before You Retrieve: Agentic Planning for Multi-modal RAG

Tianyu Yang, Shir Simon, Zhenzhen Li +2

Multimodal retrieval-augmented generation (mRAG) aims to answer image-text queries with external knowledge, but most existing systems still retrieve directly from raw multimodal in…

cs.LG2026

Getting Better at Working With You: Compiling User Corrections into Runtime Enforcement for Coding Agents

Yujun Zhou, Kehan Guo, Haomin Zhuang +8

Interactive LLM agents are becoming part of daily work, but they do not reliably become easier to work with over time: a correction remembered in one session may still be violated…

cs.AI2026

A Survey of Multimodal Mathematical Reasoning: From Perception, Alignment to Reasoning

Tianyu Yang, Sihong Wu, Yilun Zhao +6

Multimodal Mathematical Reasoning (MMR) has recently attracted increasing attention for its capability to solve mathematical problems involving both textual and visual modalities.…

cs.LG2026

Evolving Language Models without Labels: Majority Drives Selection, Novelty Promotes Variation

Yujun Zhou, Zhenwen Liang, Haolin Liu +7

Large language models (LLMs) are increasingly trained with reinforcement learning from verifiable rewards (RLVR), yet real-world deployment demands models that can self-improve wit…

cs.CY2025

The Role of Computing Resources in Publishing Foundation Model Research

Yuexing Hao, Yue Huang, Haoran Zhang +8

Cutting-edge research in Artificial Intelligence (AI) requires considerable resources, including Graphics Processing Units (GPUs), data, and human resources. In this paper, we eval…