Publications (61)
Federated Semi-Supervised Learning with Inter-Client Consistency & Disjoint Learning
Wonyong Jeong, Jaehong Yoon, Eunho Yang +1
Reliable and Responsible Foundation Models: A Comprehensive Survey
Xinyu Yang, Junlin Han, Rishi Bommasani +49
EPiC: Efficient Video Camera Control Learning with Precise Anchor-Video Guidance
Zun Wang, Jaemin Cho, Jialu Li +4
On the Soft-Subnetwork for Few-shot Class Incremental Learning
Haeyong Kang, Jaehong Yoon, Sultan Rizky Hikmawan Madjid +2
Confidence-Aware Tool Orchestration for Robust Video Understanding
Yangfan He, Yujin Choi, Jaehong Yoon
WorldMM: Dynamic Multimodal Memory Agent for Long Video Reasoning
Woongyeong Yeo, Kangsan Kim, Jaehong Yoon +1
On the Trustworthiness of Generative Foundation Models: Guideline, Assessment, and Perspective
Yue Huang, Chujie Gao, Siyuan Wu +63
Adaptive Network Sparsification with Dependent Variational Beta-Bernoulli Dropout
Juho Lee, Saehoon Kim, Jaehong Yoon +3
Glider: Global and Local Instruction-Driven Expert Router
Pingzhi Li, Prateek Yadav, Jaehong Yoon +4
STELLA: Continual Audio-Video Pre-training with Spatio-Temporal Localized Alignment
Jaewoo Lee, Jaehong Yoon, Wonjae Kim +2
Continual Learners are Incremental Model Generalizers
Jaehong Yoon, Sung Ju Hwang, Yue Cao
Representational Continuity for Unsupervised Continual Learning
Divyam Madaan, Jaehong Yoon, Yuanchun Li +2
Carpe Diem: On the Evaluation of World Knowledge in Lifelong Language Models
Yujin Kim, Jaehong Yoon, Seonghyeon Ye +4
MEXA: Towards General Multimodal Reasoning with Dynamic Multi-Expert Aggregation
Shoubin Yu, Yue Zhang, Ziyang Wang +2
CREMA: Generalizable and Efficient Video-Language Reasoning via Multimodal Modular Fusion
Shoubin Yu, Jaehong Yoon, Mohit Bansal
BECoTTA: Input-dependent Online Blending of Experts for Continual Test-time Adaptation
Daeun Lee, Jaehong Yoon, Sung Ju Hwang
Training-free Guidance in Text-to-Video Generation via Multimodal Planning and Structured Noise Initialization
Jialu Li, Shoubin Yu, Han Lin +3
Bitwidth Heterogeneous Federated Learning with Progressive Weight Dequantization
Jaehong Yoon, Geon Park, Wonyong Jeong +1
Frame Guidance: Training-Free Guidance for Frame-Level Control in Video Diffusion Models
Sangwon Jang, Taekyung Ki, Jaehyeong Jo +4
Video-Skill-CoT: Skill-based Chain-of-Thoughts for Domain-Adaptive Video Reasoning
Daeun Lee, Jaehong Yoon, Jaemin Cho +1
Self-Refining Video Sampling
Sangwon Jang, Taekyung Ki, Jaehyeong Jo +3
Are Video Reasoning Models Ready to Go Outside?
Yangfan He, Changgyu Boo, Jaehong Yoon
RACCooN: A Versatile Instructional Video Editing Framework with Auto-Generated Narratives
Jaehong Yoon, Shoubin Yu, Mohit Bansal
SAFREE: Training-Free and Adaptive Guard for Safe Text-to-Image And Video Generation
Jaehong Yoon, Shoubin Yu, Vaidehi Patil +2
Planning with Sketch-Guided Verification for Physics-Aware Video Generation
Yidong Huang, Zun Wang, Han Lin +5
VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos
Ziyang Wang, Shoubin Yu, Elias Stengel-Eskin +4
Lifelong Learning with Dynamically Expandable Networks
Jaehong Yoon, Eunho Yang, Jeongtae Lee +1
BiTAT: Neural Network Binarization with Task-dependent Aggregated Transformation
Geon Park, Jaehong Yoon, Haiyang Zhang +3
MuseBench: Benchmarking Intent-Level Audiovisual Arts Understanding in MLLMs
Yuxuan Fan, Gyusik Seo, Jing Hao +3
Personalized Subgraph Federated Learning
Jinheon Baek, Wonyong Jeong, Jiongdao Jin +2
Safe Few-Step Generation via Velocity Editing
Yujin Choi, Jaehong Yoon
VGGT-Edit: Feed-forward Native 3D Scene Editing with Residual Field Prediction
Kaixin Zhu, Yiwen Tang, Yifan Yang +9
EVEREST: Efficient Masked Video Autoencoder by Removing Redundant Spatiotemporal Tokens
Sunil Hwang, Jaehong Yoon, Youngwan Lee +1
EnvGen: Generating and Adapting Environments via LLMs for Training Embodied Agents
Abhay Zala, Jaemin Cho, Han Lin +2
Progressive Fourier Neural Representation for Sequential Video Compilation
Haeyong Kang, Jaehong Yoon, DaHyun Kim +2
Scalable and Order-robust Continual Learning with Additive Parameter Decomposition
Jaehong Yoon, Saehoon Kim, Eunho Yang +1
Multimodal Representation Learning by Alternating Unimodal Adaptation
Xiaohui Zhang, Jaehong Yoon, Mohit Bansal +1
Adapt-: Scalable Continual Multimodal Instruction Tuning via Dynamic Data Selection
Adyasha Maharana, Jaehong Yoon, Tianlong Chen +1
SELMA: Learning and Merging Skill-Specific Text-to-Image Experts with Auto-Generated Data
Jialu Li, Jaemin Cho, Yi-Lin Sung +2
PhyMotion: Structured 3D Motion Reward for Physics-Grounded Human Video Generation
Yidong Huang, Zun Wang, Han Lin +6
DreamRunner: Fine-Grained Compositional Story-to-Video Generation with Retrieval-Augmented Motion Adaptation
Zun Wang, Jialu Li, Han Lin +2
ECoFLaP: Efficient Coarse-to-Fine Layer-Wise Pruning for Vision-Language Models
Yi-Lin Sung, Jaehong Yoon, Mohit Bansal
DART: Leveraging Multi-Agent Disagreement for Tool Recruitment in Multimodal Reasoning
Nithin Sivakumaran, Justin Chih-Yao Chen, David Wan +4
EgoMemReason: A Memory-Driven Reasoning Benchmark for Long-Horizon Egocentric Video Understanding
Ziyang Wang, Yue Zhang, Shoubin Yu +6
Video-RTS: Rethinking Reinforcement Learning and Test-Time Scaling for Efficient and Enhanced Video Reasoning
Ziyang Wang, Jaehong Yoon, Shoubin Yu +3
Avatar Forcing: Real-Time Interactive Head Avatar Generation for Natural Conversation
Taekyung Ki, Sangwon Jang, Jaehyeong Jo +2
Self-Correcting Text-to-Video Generation with Misalignment Detection and Localized Refinement
Daeun Lee, Jaehong Yoon, Jaemin Cho +1
Federated Continual Learning with Weighted Inter-client Transfer
Jaehong Yoon, Wonyong Jeong, Giwoong Lee +2
RSQ: Learning from Important Tokens Leads to Better Quantized LLMs
Yi-Lin Sung, Prateek Yadav, Jialu Li +2
Hierarchy-Aware Multimodal Unlearning for Medical AI
Fengli Wu, Vaidehi Patil, Jaehong Yoon +2
Forget-free Continual Learning with Soft-Winning SubNetworks
Haeyong Kang, Jaehong Yoon, Sultan Rizky Madjid +2
Refusal Falls off a Cliff: How Safety Alignment Fails in Reasoning?
Qingyu Yin, Chak Tou Leong, Linyi Yang +7
Movie Facts and Fibs (MF): A Benchmark for Long Movie Understanding
Emmanouil Zaranis, António Farinhas, Saul Santos +28
Text-Conditioned Sampling Framework for Text-to-Image Generation with Masked Generative Models
Jaewoong Lee, Sangwon Jang, Jaehyeong Jo +5
Analyzing and Mitigating Object Hallucination in Large Vision-Language Models
Yiyang Zhou, Chenhang Cui, Jaehong Yoon +5
When and How Much to Imagine: Adaptive Test-Time Scaling with World Models for Visual Spatial Reasoning
Shoubin Yu, Yue Zhang, Zun Wang +4
Continual Learning: Forget-free Winning Subnetworks for Video Representations
Haeyong Kang, Jaehong Yoon, Sung Ju Hwang +1
Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences
Xiyao Wang, Yuhang Zhou, Xiaoyu Liu +9
AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
Zun Wang, Han Lin, Jaehong Yoon +3
Rapid Structural Pruning of Neural Networks with Set-based Task-Adaptive Meta-Pruning
Minyoung Song, Jaehong Yoon, Eunho Yang +1
Online Coreset Selection for Rehearsal-based Continual Learning
Jaehong Yoon, Divyam Madaan, Eunho Yang +1