collaborators

9 papers

cs.CV2026

When to Lock Attention: Training-Free KV Control in Video Diffusion

Tianyi Zeng, Jincheng Gao, Tianyi Wang +8

Maintaining background consistency while enhancing foreground quality remains a core challenge in video editing. Injecting full-image information often leads to background artifact…

cs.CV2025

UrbanSense:A Framework for Quantitative Analysis of Urban Streetscapes leveraging Vision Large Language Models

Jun Yin, Jing Zhong, Peilin Li +4

Urban cultures and architectural styles vary significantly across cities due to geographical, chronological, historical, and socio-political factors. Understanding these difference…

cs.CL2025

FloorPlan-DeepSeek (FPDS): A multimodal approach to floorplan generation using vector-based next room prediction

Jun Yin, Pengyu Zeng, Jing Zhong +4

In the architectural design process, floor plan generation is inherently progressive and iterative. However, existing generative models for floor plans are predominantly end-to-end…

cs.CV2025

Segment Any Architectural Facades (SAAF):An automatic segmentation model for building facades, walls and windows based on multimodal semantics guidance

Peilin Li, Jun Yin, Jing Zhong +3

In the context of the digital development of architecture, the automatic segmentation of walls and windows is a key step in improving the efficiency of building information models…

cs.AI2025

FloorplanMAE:A self-supervised framework for complete floorplan generation from partial inputs

Jun Yin, Jing Zhong, Pengyu Zeng +4

In the architectural design process, floorplan design is often a dynamic and iterative process. Architects progressively draw various parts of the floorplan according to their idea…

cs.CV2025

A Cascading Cooperative Multi-agent Framework for On-ramp Merging Control Integrating Large Language Models

Miao Zhang, Zhenlong Fang, Tianyi Wang +4

Traditional Reinforcement Learning (RL) suffers from replicating human-like behaviors, generalizing effectively in multi-agent scenarios, and overcoming inherent interpretability i…