2 citations · 6 across the 21 of their papers we have counts for
24 papers
CinematicVQA: Benchmarking Film-Grammar Reasoning in Large Vision-Language Models
Shuo Xing, Pooja Verlani, Balu Adsumilli +1
Cinematography, the craft of visual storytelling through framing, lighting, and camera operation, fundamentally shapes how audiences perceive and emotionally engage with video cont…
SUMO: Segment and Track Any Motion with Nonlinear State Space Models
Kexin Tian, Sixu Li, Keshu Wu +2
Visual Object Tracking (VOT) and Moving Object Segmentation (MOS) are two fundamental tasks in computer vision that involve both spatial and temporal object dynamics. Existing meth…
Knowledge Is Not Static: Order-Aware Hypergraph RAG for Language Models
Keshu Wu, Chenchen Kuai, Zihao Li +6
Retrieval-augmented generation (RAG) enhances large language models by grounding outputs in retrieved knowledge. However, existing RAG methods including graph- and hypergraph-based…
Knowing the Answer Isn't Enough: Fixing Reasoning Path Failures in LVLMs
Chaoyang Wang, Yangfan He, Yiyang Zhou +6
We reveal a critical yet underexplored flaw in Large Vision-Language Models (LVLMs): even when these models know the correct answer, they frequently arrive there through incorrect…
NexusFlow: Unifying Disparate Tasks under Partial Supervision via Invertible Flow Networks
Fangzhou Lin, Yuping Wang, Yuliang Guo +7
Partially Supervised Multi-Task Learning (PS-MTL) aims to leverage knowledge across tasks when annotations are incomplete. Existing approaches, however, have largely focused on the…
KANMixer: a minimal KAN-centered mixer for long-term time series forecasting
Lingyu Jiang, Dengzhe Hou, Yuping Wang +9
Long-term time series forecasting (LTSF) underpins critical applications from energy management to weather prediction, yet achieving reliable multi-step-ahead accuracy remains chal…