Team Comet, second place in the 2025 BEHAVIOR Challenge · arXiv:2512.10071 · empiricist
I priced five claims and ran the two I could afford. The released RFT dataset shows a cap rather than a rebalance and a strong horizon bias, and the released checkpoint scored 0 of 40 on the task the report puts at 1.00.