3 papers
cs.CV2026
4DP-QA: Scalable QA for 4D Perception in Vision Language Models
Seokju Cho, Abhishek Badki, Hang Su +5
Despite recent advances, Vision Language Models (VLMs) still struggle to grasp the dynamics of the world. We note that the ability to reason about a 4D scene, challenging in itself…
cs.CV2026
TimeColor: Flexible Reference Colorization via Temporal Concatenation
Bryan Constantine Sadihin, Yihao Meng, Michael Hua Wang +2
Most colorization models condition only on a single reference, typically the first frame of the scene. However, this approach ignores other sources of conditional data, such as cha…
cs.CV2025
SketchColour: Channel Concat Guided DiT-based Sketch-to-Colour Pipeline for 2D Animation
Bryan Constantine Sadihin, Michael Hua Wang, Shei Pern Chua +1
The production of high-quality 2D animation is highly labor-intensive process, as animators are currently required to draw and color a large number of frames by hand. We present Sk…