most citedA Large-Scale Multimodal Dataset and Benchmarks for Human Activity Scene Understanding and Reasoning

1 citations · 1 across the 4 of their papers we have counts for

collaborators

8 papers

cs.CV2025

UltraShape 1.0: High-Fidelity 3D Shape Generation via Scalable Geometric Refinement

Tanghui Jia, Dongyu Yan, Dehao Hao +11

In this report, we introduce UltraShape 1.0, a scalable 3D diffusion framework for high-fidelity 3D geometry generation. The proposed approach adopts a two-stage generation pipelin…

cs.CV20251 cited

A Large-Scale Multimodal Dataset and Benchmarks for Human Activity Scene Understanding and Reasoning

Siyang Jiang, Mu Yuan, Xiang Ji +12

Multimodal human action recognition (HAR) leverages complementary sensors for activity classification. Beyond recognition, recent advances in large language models (LLMs) enable de…

cs.LG2025

Latent Danger Zone: Distilling Unified Attention for Cross-Architecture Black-box Attacks

Yang Li, Chenyu Wang, Tingrui Wang +4

Black-box adversarial attacks remain challenging due to limited access to model internals. Existing methods often depend on specific network architectures or require numerous queri…

cs.CV2025

REGen: Multimodal Retrieval-Embedded Generation for Long-to-Short Video Editing

Weihan Xu, Yimeng Ma, Jingyue Huang +6

Short videos are an effective tool for promoting contents and improving knowledge accessibility. While existing extractive video summarization methods struggle to produce a coheren…

cs.CV2025

GL-PGENet: A Parameterized Generation Framework for Robust Document Image Enhancement

Zhihong Tang

Document Image Enhancement (DIE) serves as a critical component in Document AI systems, where its performance substantially determines the effectiveness of downstream tasks. To add…

cs.CV2025

SpaceVLLM: Endowing Multimodal Large Language Model with Spatio-Temporal Video Grounding Capability

Jiankang Wang, Zhihan Zhang, Zhihang Liu +4

Multimodal large language models (MLLMs) have made remarkable progress in either temporal or spatial localization. However, they struggle to perform spatio-temporal video grounding…