activity
20242026
most citedROOT: VLM based System for Indoor Scene Understanding and Beyond

1 citations · 1 across the 4 of their papers we have counts for

collaborators

10 papers

cs.CY2026

Evidence-Grounded Multi-Agent Planning Support for Urban Carbon Governance via RAG

Yuyan Huang, Haoran Li, Yifan Lu +3

Urban carbon governance requires planners to integrate heterogeneous evidence -- emission inventories, statistical yearbooks, policy texts, technical measures, and academic finding…

cs.CV2025

4DWorldBench: A Comprehensive Evaluation Framework for 3D/4D World Generation Models

Yiting Lu, Wei Luo, Peiyan Tu +8

World Generation Models are emerging as a cornerstone of next-generation multimodal intelligence systems. Unlike traditional 2D visual generation, World Models aim to construct rea…

cs.CL2025

SafeMT: Multi-turn Safety for Multimodal Language Models

Han Zhu, Juntao Dai, Jiaming Ji +8

With the widespread use of multi-modal Large Language models (MLLMs), safety issues have become a growing concern. Multi-turn dialogues, which are more common in everyday interacti…

cs.CL2025

StoryBox: Collaborative Multi-Agent Simulation for Hybrid Bottom-Up Long-Form Story Generation Using Large Language Models

Zehao Chen, Rong Pan, Haoran Li

Human writers often begin their stories with an overarching mental scene, where they envision the interactions between characters and their environment. Inspired by this creative p…

cs.CL2025

Semi-Supervised Synthetic Data Generation with Fine-Grained Relevance Control for Short Video Search Relevance Modeling

Haoran Li, Zhiming Su, Junyan Yao +6

Synthetic data is widely adopted in embedding models to ensure diversity in training data distributions across dimensions such as difficulty, length, and language. However, existin…

cs.CV2025

Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning

Bob Zhang, Haoran Li, Tao Zhang +5

Multimodal Large Language Models (MLLMs) perform well in single-image visual grounding but struggle with real-world tasks that demand cross-image reasoning and multi-modal instruct…