redis-payment-lock
Coordinate payments across service instances.
August 2026 sweep. No model fully solves this task.
RL Tasks
Long-horizon RL tasks for post-training. Tasks, not repo access: each ships in a runnable environment with a graded verifier, a hidden reference solution, and solver trajectories.
redis-payment-lock
August 2026 sweep. No model fully solves this task.
As of 27 Sep 2026
Each task recreates one real commit from a licensed repo. A reference solution scores at least 0.95, so every task is solvable, and it stays hidden from the solver. Graded checks give continuous reward from 0 to 1, including partial progress.
One real commit in, one verified task out.
Build the task
01
Private production code with its git history, scrubbed of secrets and personal data first.
02
One real commit becomes one task. Mechanical changes are dropped.
03
The repo before the commit, in a Harbor task container. The solver never sees the commit message.
04
Graded checks that run the code. Fakes only at external service boundaries.
Prove it, then ship it
05
Reference solution at least 0.95; doing nothing and two cheat probes at most 0.1; mutation and consistency checks hold.
06
Solver runs score the task from 0 to 1.
07
0.1 to 0.8 ships. Under 0.1 gets one de‑scope round, then a hard tier; over 0.8, an easy tier.
08
Exported as a tar.gz with its sha256, QC evidence and solver trajectories included.
$ cat instruction.md ## Task: Redis-backed payment-intent coordination Coordinate duplicate booking requests across service instances. Requests are equivalent only when vehicle, interval, and user match. For concurrent equivalent requests, perform at most one provider operation. If coordination ownership is lost mid-flight, do not report success. … 28 graded checks · live Redis beside the service checks run the code; fakes only at external service boundaries $ harbor run redis-payment-lock … 31 agent steps · 0.9M cumulative prompt tokens across steps (task medians) checks passed 19/28 (weighted) reward = 0.6786 best trial, August 2026 sweep $ cat trajectories/sweep.txt 3-trial avgs: opus-4.8 0.68 · gpt-5.6 0.45 · grok-4.5 0.44 · muse 0.10 spot check (1 trial): 0.00 no model fully solves it
task/ ├── instruction.md the task the solver reads ├── task.toml Harbor task config ├── environment/ Dockerfile + repo at the before-state ├── tests/ the graded suite ├── solution/ hidden reference patch ├── cheat/solve.sh known reward-forgery attempts ├── qa-evidence/ reference, do-nothing, mutation runs ├── trajectories/ per model and trial, with reward ├── manifest.json task manifest ├── qc-report.md QC report └── reproduction-kit.md reproduction kit
# Delivered as a tar.gz with its sha256.
Continuous reward from 0 to 1: weighted checks passed over weighted checks total.
Behavioral, integration, end-to-end and runtime checks weigh 4x; core and other checks 1x; setup, install, typecheck, build and lint checks do not count. A missing results file, altered test tooling, a verifier that leaves tracked files modified, or a run past the time cap scores 0.
On the five released samples, the reference solution and the renamed-internals run scored 1.00 and doing nothing scored 0.00. The runs ship in every bundle's qa-evidence/.
Sample tasks
Five tasks, five frontier model configurations: four at three trials each and one single-trial spot check. 3 of the 62 recorded trials fully solved a task in August 2026. Median runs of 25–75 agent steps (max 121).
Solver runs set each task’s band. Frontier-model results across several model families can be run for a delivered set on request.
Licensed directly from the teams that built it. Every bundle ships a file-tree manifest (sha256 per file plus a root hash) to check overlap with your corpus.
Every source repo passes an exposure checkEach task recreates a real change shipped in production code. Only tasks that solver runs score between 0.1 and 0.8 ship.
3 of 62 frontier trials fully solvedContinuous reward from 0 to 1 over graded checks that run the code. Probed for reward hacking before it ships, with the QC evidence in the bundle.
Reference ≥ 0.95 · doing nothing ≤ 0.1Early signal: post-training on our tasks moved Qwen3.8-27B from 1/74 to 4/74 on Terminal-Bench 3.0 in one paired pass@1 sweep.
+4.1pp on Terminal-Bench 3.0An autonomous pipeline mines, builds, and verifies tasks from 10,000+ licensed codebases. Supply scales with compute, not headcount.
1,000+ new tasks every monthHarbor task format: a runnable Docker environment, a graded verifier, and solver trajectories, built for your own RL or evaluation infrastructure.
Harbor task format, Docker, graded verifierAn early signal from one paired pass@1 sweep: we post-trained Qwen3.8-27B on our tasks and measured it on Terminal-Bench 3.0.
Context only, not a rank claim: GLM 5.2 + Claude Code is averaged over five trials per task on a different, substantially more capable harness; Qwen is one paired sweep of the 74-task snapshot, evaluated 18 Aug 2026.
Every source repo passes an exposure check before any task is built: not public, not a fork, not archived, not a published package. Every bundle ships a file-tree manifest (sha256 per file plus a root hash) so you can check overlap with your own corpus.
Live figures, rounded down, counting only codebases we can license today. Tokens cover source snapshots plus git history; historical duplicates are not deduplicated.
Also Kotlin, Ruby, Rust, C++, Dart, Terraform, Go, C, Solidity, Svelte, Elixir, and more. Repos hold several languages, so shares sum past 100%.
About 1.4M candidate code changes in test-bearing repos at today's corpus size, extrapolated from a measured mining run (201,000 candidates, 30,000 eligible after source-level screening). Candidates undergo environment construction and verifier QC before shipping.
Each task ships as a tar.gz with its sha256: the instruction, a Harbor task config, a runnable environment with the repo at the before-state, the graded test suite, a hidden reference patch, known reward-forgery attempts you can rerun yourself, QC evidence, solver trajectories with their rewards, a manifest, a QC report and a reproduction kit.
An early signal: in one paired Terminal-Bench 3.0 pass@1 sweep, Qwen3.8-27B improved from 1/74 to 4/74 (+4.1pp). The trained adapter was the intended variable; architecture, serving, Terminus-2 harness, benchmark parameters, and tasks stayed fixed.
Private production codebases we license directly from the teams that built them. Every source repo passes an exposure check before any task is built: not public, not a fork, not archived, not a published package. Every bundle ships a file-tree manifest (sha256 per file plus a root hash) so you can check overlap with your own corpus.
Request samples or book a call on this page. On a 30-minute call with the founders we walk through sample tasks, verifiers, QC evidence, and solver trajectories with your team, then scope an evaluation set against your target domains and languages.
The environment is the runnable container a task ships in: a Dockerfile and the repo at the before-state. The RL task is what runs inside it: an instruction, a graded verifier returning continuous reward from 0 to 1, and a hidden reference solution. Tasks are the unit we deliver, and each ships with its environment.
The pipeline produces 1,000+ new tasks every month. Custom sets are scoped on the intro call against your target domains, languages, and task length. Solver runs set each task’s band; frontier-model results across several model families can be run for a delivered set on request.
Per engagement, quoted after a scoping call, so the price follows the difficulty band and volume you need. Each task includes its runnable environment, verifier, reference solution, solver trajectories, and QC evidence. Licensing is non-exclusive by default.
Reward is continuous from 0 to 1: weighted checks passed over weighted checks total. Behavioral, integration, end-to-end and runtime checks weigh 4x; setup, install, typecheck, build and lint checks do not count; a failed core check scales the score by the share of core checks passed. A missing results file, altered test tooling, or a run past the time cap scores 0. Before a task ships, the reference solution scores at least 0.95 in a cold container, doing nothing scores at most 0.1, the checks are audited against the instruction, renaming internal variables keeps the score, removing one specified behavior lowers it, and two cheat probes (one exits without doing any work, one prints forged test results) score at most 0.1. Scores must also vary across solver runs, with the checks ordering attempts consistently by difficulty. The evidence ships with every task.
30 minutes with Arjun and Adam, the founders. We walk through sample tasks, verifiers, and model results, and scope a set for what you're training.
Loading available times…