Skip to content
arXiv cs.AI · Papers

Multi-turn RL with Structural and Performance Aware Rewards for CUDA Kernel Generation

arXiv:2607.20908v1 Announce Type: cross Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a powerful technique to enhance the reasoning capacity of LLMs for optimized code generation. However, existing RLVR approaches primarily rely on outcome-based signals such as correctness and speedup,