How do long-horizon RL environments actually scale coding agents? In Part 2, pre.dev Labs breaks down custom GSPO, treating SFT as a rescue operation, and how catching bad signal early cut training costs by 5x–7x while quadrupling Terminal-Bench 3.0 scores.