Labs
Research from pre.dev Labs: RL coding tasks for frontier AI labs, how they are verified and what training on them does.
6 posts
-
LabsRL Tasks: Where the reward went
Two tasks from a catalog of real commits rebuilt as graded RL environments. In both, the models did recognizable, competent work and scored close to nothing — and in both, the reason turned out to have almost nothing to do with whether they understood the problem.
-
LabsMost “Frontier” Coding Tasks Are Useless for RL [Part 2]
How do long-horizon RL environments actually scale coding agents? In Part 2, pre.dev Labs breaks down custom GSPO, treating SFT as a rescue operation, and how catching bad signal early cut training costs by 5x–7x while quadrupling Terminal-Bench 3.0 scores.
-
LabsMost “Frontier” Coding Tasks Are Useless for RL [Part 1]
Do long-horizon RL tasks move the needle for coding agents? pre.dev Labs trained Qwen3.8-27B on ~100 frontier tasks, boosting its Terminal-Bench 3.0 score from 1/74 to 4/74. Here is how we built the environments and scaled agent performance.
-
LabsHow We Beat Claude Opus with a Smaller Model.
Stop burning tokens on brute force. If you’re using coding agents but seeing diminished ROI, this breakdown is for you.
-
Browser Agentspre.dev Browser Agents: Our First Labs Project is 3.4x Faster and 2.3x Cheaper Than Browser-Use
Introducing pre.dev labs browser agents. We are sharing our benchmark performance & some exciting use cased form our early access customers.
-
AIFrontier labs won't build good harnesses. Their incentives won't let them.
Here's the problem nobody at Anthropic or OpenAI wants to say out loud: you can't build a token-efficient harness when your entire business model depends on burning tokens.Anthropic just committed…
Get new posts by email
New posts from pre.dev.