activity
20242026
collaborators

12 papers

cs.AI2026

Improving Generalization Robustness of Multimodal RLVR

Pengfei Zhou, Zhiwei Tang, Xiaopeng Peng +11

Reinforcement Learning with Verifiable Rewards (RLVR) makes Multimodal Large Language Models more accurate, but the gains are brittle: simply paraphrasing a question or changing th…

cs.DC2026

FeLoG: Scalable and Efficient Distributed Graph Embedding with Feedback Loop Mechanism

Peng Fang, Arijit Khan, Ziqiang Wu +4

Graph embedding maps graph nodes into low-dimensional vectors to support applications such as recommendation, fraud detection, and graph-based retrieval-augmented generation (Graph…

cs.CV2026

WorldOlympiad: Can Your World Model Survive a Triathlon?

Yuke Zhao, Wangbo Zhao, Weijie Wang +8

We introduce WorldOlympiad, a benchmark for diagnosing video-based world models across physical faithfulness, geometric consistency, and interaction fidelity. While existing benchm…

cs.CV2026

BIFE: Better Interaction, Fewer Errors for Minute-Long Video Generation

Zeyu Zhang, Jinyuan Mao, Shuning Chang +5

Long video generation is a critical step toward building realistic world models, requiring both high visual fidelity and long-range interaction consistency. Recent autoregressive d…

cs.CV2026

Towards Error-Free Long Video Generation

Shuning Chang, Weihua Chen, Jiasheng Tang +8

Recent advances in video generation have made minute-level synthesis possible; however, generating long videos remains challenging due to error accumulation, attribute drift, and t…

cs.LG2026

Graph Grounded Cross Attention Transformer Neural Network for Structurally Constrained Full Event Sequence Generation in Predictive Process Monitoring

Fang Wang, Ernesto Damiani

Structurally constrained event sequence generation remains challenging because generated paths must preserve transition feasibility, temporal order, termination, and attribute cons…