RL Tasks

Hard, verified coding tasks. From code no model has seen.

Long-horizon RL tasks for post-training. Tasks, not repo access: each ships in a runnable environment with a graded verifier, a hidden reference solution, and solver trajectories.

Sample task

redis-payment-lock

Coordinate payments across service instances.

Live Redis28 behavioral checksContinuous reward
Best recorded trial19 / 28 checks

August 2026 sweep. No model fully solves this task.

Task contents

ScalePipeline and corpus, updated every minute

As of 27 Sep 2026

Inside a task

A real task, end to end.

Each task recreates one real commit from a licensed repo. A reference solution scores at least 0.95, so every task is solvable, and it stays hidden from the solver. Graded checks give continuous reward from 0 to 1, including partial progress.

How an RL task is made

One real commit in, one verified task out.

Build the task

  1. 01

    Licensed repo

    Private production code with its git history, scrubbed of secrets and personal data first.

  2. 02

    Mine one real commit

    One real commit becomes one task. Mechanical changes are dropped.

  3. 03

    Rebuild the before-state

    The repo before the commit, in a Harbor task container. The solver never sees the commit message.

  4. 04

    Write the verifier

    Graded checks that run the code. Fakes only at external service boundaries.

Prove it, then ship it

  1. 05

    QC checks

    Reference solution at least 0.95; doing nothing and two cheat probes at most 0.1; mutation and consistency checks hold.

  2. 06

    Calibrate

    Solver runs score the task from 0 to 1.

  3. 07

    Band

    0.1 to 0.8 ships. Under 0.1 gets one de‑scope round, then a hard tier; over 0.8, an easy tier.

  4. 08

    Bundle

    Exported as a tar.gz with its sha256, QC evidence and solver trajectories included.

$ cat instruction.md
## Task: Redis-backed payment-intent coordination
Coordinate duplicate booking requests across service instances.
Requests are equivalent only when vehicle, interval, and user
match. For concurrent equivalent requests, perform at most one
provider operation. If coordination ownership is lost mid-flight,
do not report success. …
28 graded checks · live Redis beside the service
checks run the code; fakes only at external service boundaries

$ harbor run redis-payment-lock
… 31 agent steps · 0.9M cumulative prompt tokens across steps (task medians)
checks passed 19/28 (weighted)
reward = 0.6786   best trial, August 2026 sweep

$ cat trajectories/sweep.txt
3-trial avgs: opus-4.8 0.68 · gpt-5.6 0.45 · grok-4.5 0.44 · muse 0.10
spot check (1 trial): 0.00
no model fully solves it
task/
├── instruction.md       the task the solver reads
├── task.toml            Harbor task config
├── environment/         Dockerfile + repo at the before-state
├── tests/               the graded suite
├── solution/            hidden reference patch
├── cheat/solve.sh       known reward-forgery attempts
├── qa-evidence/         reference, do-nothing, mutation runs
├── trajectories/        per model and trial, with reward
├── manifest.json        task manifest
├── qc-report.md         QC report
└── reproduction-kit.md  reproduction kit

# Delivered as a tar.gz with its sha256.

The reward

Continuous reward from 0 to 1: weighted checks passed over weighted checks total.

Behavioral, integration, end-to-end and runtime checks weigh 4x; core and other checks 1x; setup, install, typecheck, build and lint checks do not count. A missing results file, altered test tooling, a verifier that leaves tracked files modified, or a run past the time cap scores 0.

QC checks every shipped task clears
  • Reference solution, cold containerat least 0.95
  • Doing nothingat most 0.1
  • Fairness auditchecks match the instruction
  • Rename internal variablesholds at 0.95 or above
  • Remove one specified behaviorscore must drop
  • Cheat probes: exit without working, forge resultsat most 0.1
  • Scores vary across solver runsnot flat
  • Checks order attempts by difficultyLoevinger H ≥ 0.40
  • Solver runs0.1 to 0.8

On the five released samples, the reference solution and the renamed-internals run scored 1.00 and doing nothing scored 0.00. The runs ship in every bundle's qa-evidence/.

Sample tasks

Frontier models score them. Few score well.

Five tasks, five frontier model configurations: four at three trials each and one single-trial spot check. 3 of the 62 recorded trials fully solved a task in August 2026. Median runs of 25–75 agent steps (max 121).

each dot: one model config's avg reward · zeros nudged apart for visibility
non-verbal-vocalization
0.13 ± 0.26
admin-panel-fixes
0.30 ± 0.23
redis-payment-lock
0.38 ± 0.23
prosody-extraction
0.66 ± 0.35
sms-verification
0.67 ± 0.25

Solver runs set each task’s band. Frontier-model results across several model families can be run for a delivered set on request.

Why pre.dev

Why frontier labs work with us.

  • Code no model has seen

    Licensed directly from the teams that built it. Every bundle ships a file-tree manifest (sha256 per file plus a root hash) to check overlap with your corpus.

    Every source repo passes an exposure check
  • Hard enough to teach

    Each task recreates a real change shipped in production code. Only tasks that solver runs score between 0.1 and 0.8 ship.

    3 of 62 frontier trials fully solved
  • Rewards you can audit

    Continuous reward from 0 to 1 over graded checks that run the code. Probed for reward hacking before it ships, with the QC evidence in the bundle.

    Reference ≥ 0.95 · doing nothing ≤ 0.1
  • Makes models better

    Early signal: post-training on our tasks moved Qwen3.8-27B from 1/74 to 4/74 on Terminal-Bench 3.0 in one paired pass@1 sweep.

    +4.1pp on Terminal-Bench 3.0
  • Volume that keeps coming

    An autonomous pipeline mines, builds, and verifies tasks from 10,000+ licensed codebases. Supply scales with compute, not headcount.

    1,000+ new tasks every month
  • Drops into your stack

    Harbor task format: a runnable Docker environment, a graded verifier, and solver trajectories, built for your own RL or evaluation infrastructure.

    Harbor task format, Docker, graded verifier
Early signal

Trained on private tasks. Measured on a public frontier bench.

An early signal from one paired pass@1 sweep: we post-trained Qwen3.8-27B on our tasks and measured it on Terminal-Bench 3.0.

Terminal-Bench 3.0, before and after
pass@1 · same architecture and serving · same Terminus-2 harness · trained adapter changes
0%2%4%6%8%Context: GLM 5.2 + Claude Code4.6% ± 1.0%1.4% · 1/745.4% · 4/74base Qwen3.8-27Bafter RL on pre.dev tasks

Context only, not a rank claim: GLM 5.2 + Claude Code is averaged over five trials per task on a different, substantially more capable harness; Qwen is one paired sweep of the 74-task snapshot, evaluated 18 Aug 2026.

Only the adapter changed
Same architecture, serving configuration, Terminus-2 harness, benchmark parameters, and 74-task set in both arms.
No overlap with the benchmark
Our repository- and task-overlap checks found no shared repositories or tasks between the training environments and Terminal-Bench 3.0.
Everything is reviewable
Trajectories, verifiers, adapters, and the full transfer report are available for review.
The corpus

10,000+ private production codebases. Licensed from the teams that built them.

Every source repo passes an exposure check before any task is built: not public, not a fork, not archived, not a published package. Every bundle ships a file-tree manifest (sha256 per file plus a root hash) so you can check overlap with your own corpus.

50B+
tokens of private code
2.5M+
commits of history
25+
languages
2008
commit history since

Live figures, rounded down, counting only codebases we can license today. Tokens cover source snapshots plus git history; historical duplicates are not deduplicated.

Languages · 25+% of repos
JavaScript62%TypeScript32%Python16%PHP12%SQL11%Java9%Swift6%Vue5%C#5%Objective-C4%

Also Kotlin, Ruby, Rust, C++, Dart, Terraform, Go, C, Solidity, Svelte, Elixir, and more. Repos hold several languages, so shares sum past 100%.

Engineering domains
Frontend44%Backend36%Full-stack12%Infra & DevOps5%Data engineering2%ML1%AI research1%Security<1%

About 1.4M candidate code changes in test-bearing repos at today's corpus size, extrapolated from a measured mining run (201,000 candidates, 30,000 eligible after source-level screening). Candidates undergo environment construction and verifier QC before shipping.

FAQ

Questions post-training teams ask.

What is included in a task?

Each task ships as a tar.gz with its sha256: the instruction, a Harbor task config, a runnable environment with the repo at the before-state, the graded test suite, a hidden reference patch, known reward-forgery attempts you can rerun yourself, QC evidence, solver trajectories with their rewards, a manifest, a QC report and a reproduction kit.

What improvement have you measured?

An early signal: in one paired Terminal-Bench 3.0 pass@1 sweep, Qwen3.8-27B improved from 1/74 to 4/74 (+4.1pp). The trained adapter was the intended variable; architecture, serving, Terminus-2 harness, benchmark parameters, and tasks stayed fixed.

Where do the codebases come from?

Private production codebases we license directly from the teams that built them. Every source repo passes an exposure check before any task is built: not public, not a fork, not archived, not a published package. Every bundle ships a file-tree manifest (sha256 per file plus a root hash) so you can check overlap with your own corpus.

How do we evaluate before buying?

Request samples or book a call on this page. On a 30-minute call with the founders we walk through sample tasks, verifiers, QC evidence, and solver trajectories with your team, then scope an evaluation set against your target domains and languages.

What is the difference between an RL environment and an RL task?

The environment is the runnable container a task ships in: a Dockerfile and the repo at the before-state. The RL task is what runs inside it: an instruction, a graded verifier returning continuous reward from 0 to 1, and a hidden reference solution. Tasks are the unit we deliver, and each ships with its environment.

How fast can you deliver a custom set of RL tasks?

The pipeline produces 1,000+ new tasks every month. Custom sets are scoped on the intro call against your target domains, languages, and task length. Solver runs set each task’s band; frontier-model results across several model families can be run for a delivered set on request.

How are RL tasks priced?

Per engagement, quoted after a scoping call, so the price follows the difficulty band and volume you need. Each task includes its runnable environment, verifier, reference solution, solver trajectories, and QC evidence. Licensing is non-exclusive by default.

How do the verifiers resist reward hacking?

Reward is continuous from 0 to 1: weighted checks passed over weighted checks total. Behavioral, integration, end-to-end and runtime checks weigh 4x; setup, install, typecheck, build and lint checks do not count; a failed core check scales the score by the share of core checks passed. A missing results file, altered test tooling, or a run past the time cap scores 0. Before a task ships, the reference solution scores at least 0.95 in a cold container, doing nothing scores at most 0.1, the checks are audited against the instruction, renaming internal variables keeps the score, removing one specified behavior lowers it, and two cheat probes (one exits without doing any work, one prints forged test results) score at most 0.1. Scores must also vary across solver runs, with the checks ordering attempts consistently by difficulty. The evidence ships with every task.

Book an intro call

30 minutes with Arjun and Adam, the founders. We walk through sample tasks, verifiers, and model results, and scope a set for what you're training.

Loading available times…