4 papers
From Final Artifacts to Trajectories: Retrospective Process Supervision for Evidence-Grounded Long-Form Generation
Junjie Huang, Jiarui Qin, Di Yin +4
Trajectory data is getting more vital for training large language models for boosting the agentic abilities. Unlike the verifiable domains such as coding or mathematics, scaling tr…
Weaving Visual Narratives: Agentic Image Bundle Composition Beyond Atomic Visual Matching
Rong Shan, Tianyi Xu, Congmin Zheng +9
Image retrieval has traditionally been formulated as a point-wise matching problem, where each candidate image is scored in isolation. However, this atomic paradigm fails to captur…
PhysGM: Large Physical Gaussian Model for Feed-Forward 4D Synthesis
Chunji Lv, Zequn Chen, Donglin Di +5
Despite advances in physics-based 3D motion synthesis, current methods face key limitations: reliance on pre-reconstructed 3D Gaussian Splatting (3DGS) built from dense multi-view…
LLaPa: A Vision-Language Model Framework for Counterfactual-Aware Procedural Planning
Shibo Sun, Xue Li, Donglin Di +6
While large language models (LLMs) have advanced procedural planning for embodied AI systems through strong reasoning abilities, the integration of multimodal inputs and counterfac…