Publications (52)
IllumiCraft: Unified Geometry and Illumination Diffusion for Controllable Video Generation
Yuanze Lin, Yi-Wen Chen, Yi-Hsuan Tsai +2
Although diffusion-based models can generate high-quality and high-resolution video sequences from textual or image inputs, they lack explicit integration of geometric cues when co…
PIXLRelight: Controllable Relighting via Intrinsic Conditioning
Miguel Farinha, Ronald Clark
We present PIXLRelight, a feed-forward approach for physically controllable single-image relighting. Existing methods either provide limited lighting control (e.g. through text or…
TermiNeRF: Ray Termination Prediction for Efficient Neural Rendering
Martin Piala, Ronald Clark
Volume rendering using neural fields has shown great promise in capturing and synthesizing novel views of 3D scenes. However, this type of approach requires querying the volume net…
Reducing Annotation Burden in Physical Activity Research Using Vision-Language Models
Abram Schonfeldt, Benjamin Maylor, Xiaofang Chen +2
Introduction: Data from wearable devices collected in free-living settings, and labelled with physical activity behaviours compatible with health research, are essential for both v…
Olympus: A Universal Task Router for Computer Vision Tasks
Yuanze Lin, Yunsheng Li, Dongdong Chen +3
We introduce Olympus, a new approach that transforms Multimodal Large Language Models (MLLMs) into a unified framework capable of handling a wide array of computer vision tasks. Ut…
Reg4Pru: Regularisation Through Random Token Routing for Token Pruning
Julian Wyatt, Ronald Clark, Irina Voiculescu
Transformers are widely adopted in modern vision models due to their strong ability to scale with dataset size and generalisability. However, this comes with a major drawback: comp…