Adaptive Rollout Allocation for Online Reinforcement Learning with Verifiable Rewards

: 10h15, ngày 17/09/2026 (Thứ Năm)

: C9-303

: NCM Xác suất thống kê và ứng dụng

: TS. Nguyễn Việt Anh

: Department of Systems Engineering and Engineering Management, The Chinese University of Hong Kong

Tóm tắt báo cáo

Sampling efficiency is a key bottleneck in reinforcement learning with verifiable rewards. Existing group-based policy optimization methods, such as GRPO, allocate a fixed number of rollouts for all training prompts. This uniform allocation implicitly treats all prompts as equally informative and could lead to inefficient computational budget usage and impede training progress. We introduce VIP, a Variance-Informed Predictive allocation strategy that allocates a given rollout budget to the prompts in the incumbent batch to minimize the expected gradient variance of the policy update. At each iteration, VIP uses a lightweight Gaussian process model to predict per-prompt success probabilities based on recent rollouts. These probability predictions are translated into variance estimates, which are then fed into a convex optimization problem to determine the optimal rollout allocations under a hard compute budget constraint. Empirical results show that VIP consistently improves sampling efficiency and achieves higher performance than uniform or heuristic allocation strategies in multiple benchmarks.

Từ khoá: học tăng cường, phần thưởng kiểm chứng được, GRPO, quá trình Gauss, tối ưu lồi, phân bổ ngân sách tính toán.

Seminar mở cho giảng viên, nghiên cứu viên, nghiên cứu sinh, học viên cao học và sinh viên quan tâm.

Link đăng ký: https://forms.cloud.microsoft/r/dZ2tm2Rxap?origin=lprLink

 Đầu mối liên hệ: email: anh.nguyenthingoc@hust.edu.vn

Thông tin chi tiết về báo cáo viên: TS. Nguyễn Việt Anh

Assistant Professor, Department of Systems Engineering and Engineering Management, The Chinese University of Hong Kong (CUHK).

Tiến sĩ Quản lý Công nghệ, EPFL (2019); nghiên cứu viên sau tiến sĩ tại Đại học Stanford (2019–2021); Research Scientist phụ trách Học máy và Học sâu tại VinAI Research (2021–2022). Cử nhân và Thạc sĩ Industrial and Systems Engineering, NUS.

Hướng nghiên cứu: ra quyết định dưới bất định ở quy mô lớn, tối ưu hoá thống kê và học máy, AI có trách nhiệm. Giải Nhất George Nicholson Student Paper Competition (INFORMS, 2018).

Website: https://vietanhnguyen.net 

Email: nguyen (at) se.cuhk.edu.hk


Đánh giá bài viết


Xem thêm