5 papers
ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents
Fanqing Meng, Lingxiao Du, Zijian Wu +46
Language-model agents are increasingly used as persistent coworkers that assist users across multiple working days. During such workflows, the surrounding environment may change in…
Accelerated Price Adjustment for Fisher Markets with Exact Recovery of Competitive Equilibrium
He Chen, Chonghe Jiang, Anthony Man-Cho So
The canonical price-adjustment process, tâtonnement, typically fails to converge to the exact competitive equilibrium (CE) and requires a high iteration complexity of $\tilde{\math…
Probe-Free Low-Rank Activation Intervention
Chonghe Jiang, Bao Nguyen, Anthony Man-Cho So +1
Language models (LMs) can produce texts that appear accurate and coherent but contain untruthful or toxic content. Inference-time interventions that edit the hidden activations hav…
Computing Competitive Equilibrium for Chores: Linear Convergence and Lightweight Iteration
He Chen, Chonghe Jiang, Anthony Man-Cho So
Competitive equilibrium (CE) for chores has recently attracted significant attention, with many algorithms proposed to approximately compute it. However, existing algorithms either…
GETS: Ensemble Temperature Scaling for Calibration in Graph Neural Networks
Dingyi Zhuang, Chonghe Jiang, Yunhan Zheng +2
Graph Neural Networks deliver strong classification results but often suffer from poor calibration performance, leading to overconfidence or underconfidence. This is particularly p…