RLHF
Publications 3
Looking Through the Mirror: Minimax-Optimal Regularized Regrets in Online Learning and Bandits New
NeurIPS 2026
CKAIA 2026
· Distinguished Paper Award
Provably Efficient Regularized Online RLHF with Generalized Bilinear Preferences New
NeurIPS 2026
ICML 2026 Pluralistic Alignment Workshop
CKAIA 2026
· Distinguished Paper Award
Pointwise or Pairwise: When Do Pairwise Losses Help Reward Learning, Provably?
arXiv:2609.37209