Beyond Two-Stage Training: Cooperative SFT and RL for LLM Reasoning

arXiv CS Tuesday 02 June 2026, 04:00 UTC By Liang Chen, Xueting Han, Li Shen, Jing Bai, Kam-Fai Wong 1 min read

Key Points

arXiv:2509.06948v3 Announce Type: replace Abstract: Supervised fine-tuning (SFT) and reinforcement learning with verifiable rewards (RLVR) are two widely used post-training paradigms for improving the reasoning ability of large language models (LLMs). Recent methods attempt to integrate SFT and RLVR in a single stage by reweighting or scheduling their objectives. However, such coupling can be counterproductive because supervised updates are not uniformly beneficial for reward optimization. To address this, we propose BRIDGE, a scalable framework in which SFT learns to supervise RL by selectively transferring knowledge that improves reward optimization. Specifically, BRIDGE alternates two updates at each meta-training step: a base-model update that fuses the SFT and RL gradients, and an update to a lightweight low-rank adapter (LoRA) that coordinates the two objectives by maximizing a cooperative-gain signal, defined as the reward of joint SFT-RL training over an RL-only baseline. Across five mathematical reasoning benchmarks, BRIDGE consistently outperforms two-stage cold start, naive mixing, and representative single-stage integration baselines, yielding over three points average absolute improvement and more stable training dynamics. We further show that BRIDGE extends to logical reasoning and generalizes out-of-distribution to code and science without additional training, while staying robust under noisy rewards.

Cooperative SFT (ORG) RL (ORG) SFT (ORG) SFT-RL (ORG)

Originally published by arXiv CS Read original →

Beyond Two-Stage Training: Cooperative SFT and RL for LLM Reasoning

Related Stories

Could the next Chinese threat walk into your kitchen on two battery-powered legs?

Here's why universal basic income would be a disaster for America’s future

Waymo built a virtual driver to study how humans react to surprises on the road

WhatsApp ordered to host rival AI assistants for free