activity
20242026
collaborators

6 papers

cs.LG2026

MODE: Modality-Decomposed Expert-Level Mixed-Precision Quantization for MoE Multimodal LLMs

Yuanteng Chen, Peisong Wang, Zhilei Liu +9

Mixture-of-Experts Multimodal Large Language Models (MoE-MLLMs) offer remarkable performance but incur prohibitive GPU memory costs, making compression essential. Among PTQ methods…

cs.CV2026

LongSpace: Exploring Long-Horizon Spatial Memory from Perception to Recall in Video

Shiqiang Lang, Jing Liu, Haoyang He +6

Multimodal Large Language Models (MLLMs) have advanced image and video understanding and can increasingly handle longer visual inputs. Long-horizon tasks such as autonomous driving…

cs.CV2026

V2V-Bench: A Comprehensive Benchmark for Video-to-Video Generation Evaluation

Tao Liu, Leela Krishna, Gouti Pavan Kumar +2

Video-to-video (V2V) generation is difficult to evaluate because outputs must both follow editing instructions and preserve frame-level correspondence with the source video, which…

cs.AI2025

GUI-Rise: Structured Reasoning and History Summarization for GUI Navigation

Tao Liu, Chongyu Wang, Rongjie Li +3

While Multimodal Large Language Models (MLLMs) have advanced GUI navigation agents, current approaches face limitations in cross-domain generalization and effective history utiliza…

cs.CV2025

Relation-aware Hierarchical Prompt for Open-vocabulary Scene Graph Generation

Tao Liu, Rongjie Li, Chongyu Wang +1

Open-vocabulary Scene Graph Generation (OV-SGG) overcomes the limitations of the closed-set assumption by aligning visual relationship representations with open-vocabulary textual…

cs.RO2024

FastGrasp: Efficient Grasp Synthesis with Diffusion

Xiaofei Wu, Tao Liu, Caoji Li +3

Effectively modeling the interaction between human hands and objects is challenging due to the complex physical constraints and the requirement for high generation efficiency in ap…