pre.dev
  • Home
  • Blog
  • Pricing
  • Sign Up
Sign in Subscribe

Labs

A collection of 4 posts
Most “Frontier” Coding Tasks Are Useless for RL Cover Image
Labs

Most “Frontier” Coding Tasks Are Useless for RL [Part 2]

How do long-horizon RL environments actually scale coding agents? In Part 2, pre.dev Labs breaks down custom GSPO, treating SFT as a rescue operation, and how catching bad signal early cut training costs by 5x–7x while quadrupling Terminal-Bench 3.0 scores.
21 Aug 2026 5 min read
Most “Frontier” Coding Tasks Are Useless for RL Cover Image
Labs

Most “Frontier” Coding Tasks Are Useless for RL [Part 1]

Do long-horizon RL tasks move the needle for coding agents? pre.dev Labs trained Qwen3.8-27B on ~100 frontier tasks, boosting its Terminal-Bench 3.0 score from 1/74 to 4/74. Here is how we built the environments and scaled agent performance.
20 Aug 2026 7 min read
TerminalBench2.0 pre.dev (56.2%) vs Claude Code (53.9%)
Labs

How We Beat Claude Opus with a Smaller Model.

Stop burning tokens on brute force. If you’re using coding agents but seeing diminished ROI, this breakdown is for you.
20 May 2026 2 min read
pre.dev browser agents benchmark highlights (100/100 pass rate, 3.4x faster, 2.3x cheaper)
Browser Agents

pre.dev Browser Agents: Our First Labs Project is 3.4x Faster and 2.3x Cheaper Than Browser-Use

Introducing pre.dev labs browser agents. We are sharing our benchmark performance & some exciting use cased form our early access customers.
11 May 2026 7 min read
Page 1 of 1
pre.dev © 2026
  • Sign Up
Powered by Ghost