6 papers
ATHENA: Accelerated Multi-Task Heterogeneous Influence Functions for Robot Data Curation
Tao Xu, Jiaxin Wang, Runhao Zhang +7
In robot imitation learning, influence functions provide a principled approach to quantify each demonstration's effect on robot task outcomes, yet scaling them to billion-parameter…
ToolPRM: Fine-Grained Inference Scaling of Structured Outputs for Function Calling
Jianghao Lin, Yuanyuan Shi, Xin Peng +10
Large language models (LLMs) excel at function calling, but inference scaling has been explored mainly for unstructured generation. We propose an inference-scaling framework for st…
L1 Sample Flow for Efficient Visuomotor Learning
Weixi Song, Zhetao Chen, Tao Xu +6
Denoising-based models, such as diffusion and flow matching, have been a critical component of robotic manipulation for their strong distribution-fitting and scaling capacity. Conc…
AstroMMBench: A Benchmark for Evaluating Multimodal Large Language Models Capabilities in Astronomy
Jinghang Shi, Xiaoyu Tang, Yang Huang +4
Astronomical image interpretation presents a significant challenge for applying multimodal large language models (MLLMs) to specialized scientific tasks. Existing benchmarks focus…
High Performance Space Debris Tracking in Complex Skylight Backgrounds with a Large-Scale Dataset
Guohang Zhuang, Weixi Song, Jinyang Huang +3
With the rapid development of space exploration, space debris has attracted more attention due to its potential extreme threat, leading to the need for real-time and accurate debri…
Kimi-Audio Technical Report
KimiTeam, Ding Ding, Zeqian Ju +37
We present Kimi-Audio, an open-source audio foundation model that excels in audio understanding, generation, and conversation. We detail the practices in building Kimi-Audio, inclu…