activity
20232026
most citedWorldGPT: A Sora-Inspired Video AI Agent as Rich World Models from Text and Image Inputs

4 citations · 9 across the 7 of their papers we have counts for

collaborators

8 papers

cs.AI2026

Are Prompt Optimizers Blind? Cross-Modal Visual Feedback for Automatic Prompt Optimization

Haoyue Liu, Xiaoyu Ma, Ye Chen +2

Automatic prompt optimization (APO) has been widely adopted to adapt vision-language models (VLMs) to downstream tasks without weight updates, yielding promising results. However,…

cs.CL2026

Thinking by Subtraction: Confidence-Driven Contrastive Decoding for LLM Reasoning

Lexiang Tang, Weihao Gao, Bingchen Zhao +4

Recent work on test-time scaling for large language model (LLM) reasoning typically assumes that allocating more inference-time computation uniformly improves correctness. However,…

cs.CV2025

Not All Tokens and Heads Are Equally Important: Dual-Level Attention Intervention for Hallucination Mitigation

Lexiang Tang, Xianwei Zhuang, Bang Yang +5

Large vision-language models (LVLMs) have demonstrated impressive capabilities across diverse multimodal tasks, yet they remain highly susceptible to visual hallucinations (VH), of…

cs.CV20241 cited

VisionGPT-3D: A Generalized Multimodal Agent for Enhanced 3D Vision Understanding

Chris Kelly, Luhui Hu, Jiayin Hu +7

The evolution of text to visual components facilitates people's daily lives, such as generating image, videos from text and identifying the desired elements within the images. Comp…

cs.CV20243 cited

VisionGPT: Vision-Language Understanding Agent Using Generalized Multimodal Framework

Chris Kelly, Luhui Hu, Bang Yang +7

With the emergence of large language models (LLMs) and vision foundation models, how to combine the intelligence and capacity of these open-sourced or API-available models to achie…

cs.CV20244 cited

WorldGPT: A Sora-Inspired Video AI Agent as Rich World Models from Text and Image Inputs

Deshun Yang, Luhui Hu, Yu Tian +5

Several text-to-video diffusion models have demonstrated commendable capabilities in synthesizing high-quality video content. However, it remains a formidable challenge pertaining…