8 citations · 43 across the 96 of their papers we have counts for
29 papers · 1 filter
Accelerated Decoding of Centroid Positional Encoding for Instance Segmentation
Carmelo Scribano, Filippo Muzzini, Nedyalko Prisadnikov +6
Beyond model inference, the decoding stage, which converts raw network outputs into task-level representations, constitutes a significant portion of the execution cost. Despite its…
SceneBench: A Hierarchical Benchmark for Vision-Language Understanding of 3D Scenes
Anubhav Khanal, Prabigya Acharya, Roshni Poudel +5
Vision-language models excel at 2D image understanding but remain limited in 3D spatial reasoning. Progress is hindered by limitations in current benchmarks. First, 3D datasets oft…
ConeGaussian: Anti-Aliased Gaussian Ray-Tracing for Generic Central Cameras
Deheng Zhang, Letian Shi, Runyi Yang +6
In rendering, a camera is a sampling operator that maps each finite pixel to a bundle of rays. Different camera models change the geometry of this bundle, thus making a unified and…
LangStreet: Persistent Language Fields for Anchor-Decoded Street Gaussians
Runyi Yang, Deheng Zhang, Xiaoye Wang +6
Language Gaussian fields implicitly assume that the primitive carrying semantics remains identifiable across views. This assumption breaks in scalable anchor-decoded representation…
Kairos: A Dataset for Fine-Grained Video-Language Modeling over Space, Time, and Dynamics
Ruibo Ming, Lei Sun, Deheng Zhang +8
Many emerging video language modeling tasks require systems to move beyond clip-level abstraction and model visual content as it unfolds over extended time horizons. However, most…
FRAMEWORKERS: A Dynamic Multi-Agent Framework for AI-Generated Video Production
Zhendong Li, Lei Sun, Letian Shi +8
Modern video generators excel at synthesizing individual clips, but complete video production requires coordinating a long sequence of interdependent creative steps, including scri…