RL Tasks
Verified RL coding tasks built from private production code, sold to frontier AI labs for model post-training: how they are made, graded and what models score.
3 posts
-
LabsRL Tasks: Where the reward went
Two tasks from a catalog of real commits rebuilt as graded RL environments. In both, the models did recognizable, competent work and scored close to nothing — and in both, the reason turned out to have almost nothing to do with whether they understood the problem.
-
LabsMost “Frontier” Coding Tasks Are Useless for RL [Part 2]
How do long-horizon RL environments actually scale coding agents? In Part 2, pre.dev Labs breaks down custom GSPO, treating SFT as a rescue operation, and how catching bad signal early cut training costs by 5x–7x while quadrupling Terminal-Bench 3.0 scores.
-
LabsMost “Frontier” Coding Tasks Are Useless for RL [Part 1]
Do long-horizon RL tasks move the needle for coding agents? pre.dev Labs trained Qwen3.8-27B on ~100 frontier tasks, boosting its Terminal-Bench 3.0 score from 1/74 to 4/74. Here is how we built the environments and scaled agent performance.
Get new posts by email
New posts from pre.dev.