9 papers
Multi-Turn Agentic Scientific Literature Search via Workflow Induction
Jisen Li, Bingxuan Li, Nanyi Jiang +10
Scientific literature search often requires more than retrieving papers from a single query: users' intents are underspecified, preference-dependent, and evolve through interaction…
MoVA: Learning Asymmetric Dual Projections for Modular Long Video-Text Alignment
Peiyuan Zhu, Shaoan Xie, Zijian Li +5
Contrastive pre-training has propelled video-text alignment, yet models often inherit the critical limitations of their image-text predecessors like CLIP, resulting in entangled re…
On the Identification of Temporally Causal Representation with Instantaneous Dependence
Zijian Li, Yifan Shen, Kaitao Zheng +5
Temporally causal representation learning aims to identify the latent causal process from time series observations, but most methods require the assumption that the latent causal p…
Towards Identifiability of Hierarchical Temporal Causal Representation Learning
Zijian Li, Minghao Fu, Junxian Huang +5
Modeling hierarchical latent dynamics behind time series data is critical for capturing temporal dependencies across multiple levels of abstraction in real-world tasks. However, ex…
CausalVerse: Benchmarking Causal Representation Learning with Configurable High-Fidelity Simulations
Guangyi Chen, Yunlong Deng, Peiyuan Zhu +4
Causal Representation Learning (CRL) aims to uncover the data-generating process and identify the underlying causal variables and relations, whose evaluation remains inherently cha…
Controllable Video Generation with Provable Disentanglement
Yifan Shen, Peiyuan Zhu, Zijian Li +6
Controllable video generation remains a significant challenge, despite recent advances in generating high-quality and consistent videos. Most existing methods for controlling video…