Erle Zhu et al. Data Efficient RLVR via Off-Policy Influence
Guidance. ACL 2026. https://aclanthology.org/2026.acl-long.2141/
Shiye Lei, Zhihao Cheng, and Dacheng Tao. A Step Back: Prefix
Importance Ratio Stabilizes Policy Optimization. https://arxiv.org/abs/2601.22718
Yuheng Zhang et al. Rethinking Importance Sampling in LLM Policy
Optimization: A Cumulative Token Perspective. https://arxiv.org/abs/2605.07331
Deokgyu Yoon et al. Multi-Step Likelihood-Ratio Correction for
Reinforcement Learning with Verifiable Rewards. https://arxiv.org/abs/2605.20865
Lei Pang, Jun Luo, and Ruinan Jin. TIC-GRPO: Provable and
Efficient Optimization for Reinforcement Learning from Human Feedback.
https://arxiv.org/abs/2508.02833
Ningyuan Yang et al. GradAlign: Gradient-Aligned Data Selection
for LLM Reinforcement Learning. https://arxiv.org/abs/2602.21492
Zichao Yu et al. Mismatch Matters: On-Policy Distillation Beyond
Token Agreement. https://arxiv.org/abs/2608.09836
Shikun Li et al. LearnAlign: Reasoning Data Selection for
Reinforcement Learning in Large Language Models Based on Improved
Gradient Alignment. https://arxiv.org/abs/2506.11480
Xinyu Tang et al. Towards High Data Efficiency in Reinforcement
Learning with Verifiable Reward. ICLR 2026. https://openreview.net/forum?id=sruA4AZmZI
Hao Yi et al. Learn More with Less: Uncertainty Consistency
Guided Query Selection for RLVR. ICLR 2026. https://openreview.net/forum?id=OOTokVgBY6
Yuhan Li, Mingxu Zhang, Dazhong Shen, and Ying Sun. IRDS:
Interpretable RLVR Data Selection via Verifier-Coupled Sparse
Autoencoder Coverage. https://arxiv.org/abs/2605.28247
Jianghao Wu et al. Single-Rollout Hidden-State Dynamics for
Training-Free RLVR Data Selection. ICML 2026. https://arxiv.org/abs/2605.28631
Xinyu Zhou et al. Efficient RLVR Training via Weighted Mutual
Information Data Selection. https://arxiv.org/abs/2603.01907
Andrei Baroian and Rutger Berger. Prompt Replay: Speeding Up GRPO
with On-Policy Reuse of High-Signal Prompts. https://arxiv.org/abs/2603.21177
Hieu Trung Nguyen et al. Adaptive Rollout Allocation for Online
Reinforcement Learning with Verifiable Rewards. https://arxiv.org/abs/2602.01601
Heyang Jiang, Henry Liu, and Baharan Mirzasoleiman. Learning as
Reasoning Unfolds: Progressive Rollout Allocation for Efficient
Reinforcement Learning. https://arxiv.org/abs/2607.22002
Yu Li et al. Turning Off-Policy Tokens On-Policy: A Plug-in
Approach for Improving LLM Alignment. https://arxiv.org/abs/2607.04728
Mark Rowland et al. Conditional Importance Sampling for
Off-Policy Learning. AISTATS 2020. https://proceedings.mlr.press/v108/rowland20b.html
Yao Liu, Adith Swaminathan, Alekh Agarwal, and Emma Brunskill.
Off-Policy Policy Gradient with Stationary Distribution Correction. UAI
2020. https://proceedings.mlr.press/v115/liu20a.html
Jiawei Huang and Nan Jiang. From Importance Sampling to Doubly
Robust Policy Gradient. ICML 2020. https://proceedings.mlr.press/v119/huang20b.html
Emilie Kaufmann and Shivaram Kalyanakrishnan. Information
Complexity in Bandit Subset Selection. COLT 2013. https://proceedings.mlr.press/v30/Kaufmann13.html
Charles Spearman. Correlation Calculated from Faulty Data.
British Journal of Psychology, 1910.
William Brown. Some Experimental Results in the Correlation of
Mental Abilities. British Journal of Psychology, 1910.
Lutz Oettershagen. Top-\(k\) on
a Budget: Adaptive Ranking with Weak and Strong Oracles. https://arxiv.org/abs/2601.20989
Kazuki Nakayashiki and Keisuke Watanabe. Floor, Ceiling, and the
Fusion Gap: How Much of Crowd Reading Attention Can Machines Predict? https://arxiv.org/abs/2608.01704
Jiachen T. Wang et al. Rethinking Data Shapley for Data Selection
Tasks: Misleads and Merits. ICML 2024. https://proceedings.mlr.press/v235/wang24cg.html
Xiao Tian et al. Is Data Shapley Not Better than Random in Data
Selection? Ask NASH. https://arxiv.org/abs/2605.10684
An Yang et al. Qwen2.5 Technical Report. https://arxiv.org/abs/2412.15115
Karl Cobbe et al. Training Verifiers to Solve Math Word Problems.
https://arxiv.org/abs/2110.14168
Dan Hendrycks et al. Measuring Mathematical Problem Solving With
the MATH Dataset. https://arxiv.org/abs/2103.03874
Qiying Yu et al. DAPO: An Open-Source LLM Reinforcement Learning
System at Scale. https://arxiv.org/abs/2503.14476
Edward J. Hu et al. LoRA: Low-Rank Adaptation of Large Language
Models. https://arxiv.org/abs/2106.09685
Zhihong Shao et al. DeepSeekMath: Pushing the Limits of
Mathematical Reasoning in Open Language Models. https://arxiv.org/abs/2402.03300
Moses Charikar, Kevin Chen, and Martin Farach-Colton. Finding
Frequent Items in Data Streams. ICALP 2002.
Wenhao Zhang et al. Reusing Rollouts under Policy Lag:
Prefix-Normalized Policy Optimization for LLM Reinforcement Learning. https://arxiv.org/abs/2608.01418
Yixiu Mao, Yun Qu, Qi Wang, Heming Zou, and Xiangyang Ji. RLVR
without Ineffective Samples: Group Prioritized Off-Policy Optimization
for LLM Reasoning. https://arxiv.org/abs/2606.01281
Zhong Guan et al. Missing Old Logits in Asynchronous Agentic RL:
Semantic Mismatch and Repair Methods for Off-Policy Correction. https://arxiv.org/abs/2605.12070
Giyeong Oh, Junghyun Lee, Jaehyun Park, Youngjae Yu, Wonho Bae,
and Junhyug Noh. Random Is Hard to Beat: Active Selection in online DPO
with Modern LLMs. https://arxiv.org/abs/2604.02766
Minghao Tian, Yunfei Xie, and Chen Wei. How Off-Policy Can GRPO
Be? Mu-GRPO for Efficient LLM Reinforcement Learning. https://arxiv.org/abs/2605.17570
Haizhong Zheng, Jiawei Zhao, and Beidi Chen. Prosperity before
Collapse: How Far Can Off-Policy RL Reach with Stale Data on LLMs? https://arxiv.org/abs/2510.01161
Shuang Liang, Haoyang Zhou, Yifan Gong, Guowei Wang, and Xiting
Wang. LEEPS: Latent-Guided Explore-Exploit Prompt Sampling for Efficient
RLVR in Large Language Models. https://arxiv.org/abs/2607.28077
Dohyung Kim, Minbeom Kim, Jeonghye Kim, Sangmook Lee, Sojeong
Rhee, and Kyomin Jung. Beyond Normalization: Rethinking the Partition
Function as a Difficulty Scheduler for RLVR. https://arxiv.org/abs/2602.12642
Fredrik Cumlin. Rho-Perfect: Correlation Ceiling For Subjective
Evaluation Datasets. https://arxiv.org/abs/2602.08552
Or Zuk, Liat Ein-Dor, and Eytan Domany. Ranking Under
Uncertainty. UAI 2007, pp. 466–473. https://arxiv.org/abs/1206.5280