Learning from Teacher Continuations at Student States Paper • 2609.36246 • Published 10 days ago • 40
Selecting Diverse SFT Traces Improves Post-RL Generalization Paper • 2609.33780 • Published 11 days ago • 38
Unlocking Lossless Speedups in LLMs via Discrete Diffusion Paper • 2609.04010 • Published Sep 3 • 115
RubricEM: Meta-RL with Rubric-guided Policy Decomposition beyond Verifiable Rewards Paper • 2605.10899 • Published May 11 • 78 • 3
RubricEM: Meta-RL with Rubric-guided Policy Decomposition beyond Verifiable Rewards Paper • 2605.10899 • Published May 11 • 78