9 papers
When to Lock Attention: Training-Free KV Control in Video Diffusion
Tianyi Zeng, Jincheng Gao, Tianyi Wang +8
Maintaining background consistency while enhancing foreground quality remains a core challenge in video editing. Injecting full-image information often leads to background artifact…
UrbanSense:A Framework for Quantitative Analysis of Urban Streetscapes leveraging Vision Large Language Models
Jun Yin, Jing Zhong, Peilin Li +4
Urban cultures and architectural styles vary significantly across cities due to geographical, chronological, historical, and socio-political factors. Understanding these difference…
FloorPlan-DeepSeek (FPDS): A multimodal approach to floorplan generation using vector-based next room prediction
Jun Yin, Pengyu Zeng, Jing Zhong +4
In the architectural design process, floor plan generation is inherently progressive and iterative. However, existing generative models for floor plans are predominantly end-to-end…
Segment Any Architectural Facades (SAAF):An automatic segmentation model for building facades, walls and windows based on multimodal semantics guidance
Peilin Li, Jun Yin, Jing Zhong +3
In the context of the digital development of architecture, the automatic segmentation of walls and windows is a key step in improving the efficiency of building information models…
FloorplanMAE:A self-supervised framework for complete floorplan generation from partial inputs
Jun Yin, Jing Zhong, Pengyu Zeng +4
In the architectural design process, floorplan design is often a dynamic and iterative process. Architects progressively draw various parts of the floorplan according to their idea…
A Cascading Cooperative Multi-agent Framework for On-ramp Merging Control Integrating Large Language Models
Miao Zhang, Zhenlong Fang, Tianyi Wang +4
Traditional Reinforcement Learning (RL) suffers from replicating human-like behaviors, generalizing effectively in multi-agent scenarios, and overcoming inherent interpretability i…