From the 1 of 6 linked papers with an AI index.
6 papers
AI's Capability in Assisting Scientific Research in Physics, Astrophysics, and Cosmology II: Project Planning and Proposal Evaluation
Jia Liu, Veena Krishnaraj, Kateryna Vovk +18
The paper evaluates the ability of current large language models to generate one‑page research project plans in physics, astrophysics, and cosmology, and compares how human reviewe…
AIvilization v0: Toward Large-Scale Artificial Social Simulation with a Unified Agent Architecture and Adaptive Agent Profiles
Wenkai Fan, Shurui Zhang, Xiaolong Wang +7
AIvilization v0 is a publicly deployed large-scale artificial society that couples a resource-constrained sandbox with a unified LLM-agent architecture, aiming to sustain long-hori…
What Would an LLM Do? Evaluating Large Language Models for Policymaking to Alleviate Homelessness
Pierre Le Coz, Jia An Liu, Debarun Bhattacharjya +2
Large language models (LLMs) are increasingly being adopted in high-stakes domains. Their potential to encode evolving social contexts and to generate plausible scenarios position…
ETTRL: Balancing Exploration and Exploitation in LLM Test-Time Reinforcement Learning Via Entropy Mechanism
Jia Liu, ChangYi He, YingQiao Lin +3
Recent advancements in Large Language Models have yielded significant improvements in complex reasoning tasks such as mathematics and programming. However, these models remain heav…
Evaluating Temporal Plasticity in Foundation Time Series Models for Incremental Fine-tuning
Jia Liu, Cheng Jinguo, Xia Fang +2
Time series foundation models excel at diverse time series forecasting tasks, but their capacity for continuous improvement through incremental learning remains unexplored. We pres…
In-Dataset Trajectory Return Regularization for Offline Preference-based Reinforcement Learning
Songjun Tu, Jingbo Sun, Qichao Zhang +4
Offline preference-based reinforcement learning (PbRL) typically operates in two phases: first, use human preferences to learn a reward model and annotate rewards for a reward-free…