OpenRSI Index

Preview v0.1

What if agents take over thousands of GPUs for model training?

Signature task samples Pre-trainingLEARNING FROM SCRATCH Marin-Scaling-Ladder550M → 2.545B · E0–E5 18,432 H100-HOURS / PER RUN Post-trainingBEYOND THE PRODUCTION MODEL Qwen-122B-RL-MergeSYNTHETIC RL TASKS FOR 5 DOMAINS 20,864 H100-HOURS / PER RUN Vision-GenGENERATING IMAGES GPIC Leaderboard100M TRAINING IMAGES 5,815 H100-HOURS / PER RUN
  1. Pre-training · Signature taskMarin-Scaling-Ladder18,432 H100-hours / per run550M → 2.545B · E0–E5
  2. Post-training · Signature taskQwen-122B-RL-Merge20,864 H100-hours / per runSynthetic RL tasks for 5 domains
  3. Vision-Gen · Signature taskGPIC Leaderboard5,815 H100-hours / per run100M training images
  4. more domains to come · any stage of foundation-model development

We target tasks related to foundation-model development, with also vertical domains included.

OpenRSI Index

We're introducing OpenRSI Index. It turns influential open model-development projects into autoresearch environments where agents iterate against hidden verifiers. An agent takes over the computing cluster and optimizes a real open-source research codebase against the original human baselines.

We target to cover every stack of foundation-model development, from pre-training and data to systems, inference, and the harness, plus vertical domains such as physics, medicine, 3D vision, and science.

RSI: the last puzzle to ASI

To take ASI seriously is to accept a weak-to-strong premise: human intelligence can build a process that eventually produces intelligence beyond itself. Whether the RSI loop can move past the best human-designed method is the last puzzle on the road to ASI.

Scaling on insights. Progress is limited not only by compute and data, but by the bandwidth of insights. An agent scales both the generation of ideas and their implementation, and sweeps a far larger region of method-space. That is the ambition of RSI, and exactly the thing worth measuring carefully.

OpenRSI: Make RSI Benefit Everyone

From the community, for the community. RSI is built from generations of ideas, code, data, and experience from researchers, independent labs, and domain teams alike, and it should return the benefits to everyone who builds models. We aim to keep RSI open through shared platforms and tools, so more people can participate and benefit.

Shape RSI together. We set RSI’s standards together with the community and challenge frontier models with the hardest problems in our fields. Through about 1 hour of conversation, RSI-Anything (our contribution pipeline) helps clarify research questions, constraints, and evaluations, then packages into runnable autoresearch benchmark tasks. We want researchers and agents to solve these problems together, sharing new insights, methods, and results with the community.

Research will never end. A game turns zero-sum only when the pie is too small to share or a field has a finite ceiling. However, research is definitely the field with the highest ceiling there is. Better Human–AI collaboration makes it more prosperous to be a positive-sum game. What shifts is the mindset: it frees researchers to find and formulate the crazier, more valuable, more exciting problems in the world.

Our design principle

  1. 01Real-world research, not fabricated toys.

    Open research community will have the benchmark of our own. Every RSI environment is sourced from a real, fully open-source model research project.

  2. 02Optimization, not reproduction.

    The agent starts from the original code, data, and checkpoints to deliver an artifact that compares with the original recipe.

  3. 03Scientific discovery, not parameter sweeping.

    Form hypotheses, change methods, learn from experiments, and land gains that hold under a fixed scientific contract.

  4. 04Production scale, not proof-of-concept.

    Problems frontier researchers care about, where improvements matter and transfer to real training.

InputOpen-source project
Packaged asAgentic autoresearch environment
ProducingOpenRSI Index

Partnership

University of
Washington
UC Berkeley
University of Washington
UC Berkeley
Stanford University
Princeton University
Yale University
Massachusetts Institute of Technology
IBM Research
Texas A&M
University
University of Notre Dame
University of Rochester
Michigan State University
Amazon

OpenRSI Index is co-led with the MIT-IBM Watson AI Lab and Amazon A-EVO Lab. It is an open academic research initiative built for and with the research community.

Advisors (in alphabetical order)

Signature task samples

These samples are for an initial preview only and may be further adjusted. More samples are ongoing.

Pre-trainingOptimizer design18,432 H100-hours / per run

RungModelParamsTokensScoring updateGPUsGPU-h
E0d1152-L12550M2.904B44,317817.6
E1d1408-L15837M3.613B55,125827.6
E2d1536-L16998M4.983B38,014831.0
E3d1792-L181.385B10.560B40,28332302.5
E4d2048-L211.935B14.805B30,00032286.2
E5d2304-L232.545B18.617B35,0001282,932.1
Improvement over AdamH, per ladder rungGeometric mean of two ratios: Paloma bits-per-byte and log PPL.Each rung is a separate model, 550M to 2.5B. Higher is better.-1%-0.5%+0.5%+1%0AdamH control−1%: a rung below this fails the submissionCodex · GPT-5.6 — PSPR, +0.60% on averageClaude Code · Opus 5 — RMBT, +0.35% on averageE0550ME1837ME2998ME31.39BE41.94BE52.54BLadder rung and model sizeBetter than AdamH (%)
Best six-rung improvement during the search● new best · × rejected / not retained · ○ trial · ◇ partial result · ▪ stopped0+0.2%+0.4%+0.6%01020304050607080Cumulative runtime (hours)Best complete six-rung improvement (%)Codex · GPT-5.6 · FCPR · 3.60 h Full ladder launched Best complete six-rung improvement at this time: +0.00% 3.6h 'FCPR passes its second preregistered screen' → full ladder frozenCodex · GPT-5.6 · TARF · 10.90 h Plan not launched Best complete six-rung improvement at this time: +0.00% 9.2h preregistered contingent on T-FCPR passing; T-FCPR rejected 10.9h → never implemented or launched Nearby events on the same curve: Codex · GPT-5.6 · T-FCPR · 10.91 h Rejected / failed Best complete six-rung improvement at this time: +0.00% Rejected at E2, update 5,000: geometric gain 0.99955754 versus FCPR, despite passing the macro gate. E0 had narrowly passed at 1.00010275. Trajectory steps: 1527, 1679, 1680Codex · GPT-5.6 · FCPR · 12.30 h Paused Best complete six-rung improvement at this time: +0.00% 12.3h 'FCPR E0 25k comparison is negative … the pause is operational, not a substituted final verdict'; never resumed, never formally rejected Nearby events on the same curve: Codex · GPT-5.6 · TA-FCPR · 12.22 h Rejected / failed Best complete six-rung improvement at this time: +0.00% Rejected at both 5,000-update screens versus FCPR: geometric gain 0.98204696 on E0 and 0.98710723 on E2; both macro safety gates fail. Trajectory steps: 1852, 1856, 1912, 1915Codex · GPT-5.6 · OEPR · 15.40 h Full ladder launched Best complete six-rung improvement at this time: +0.00% 15.4h 'OEPR is promoted to a frozen full ladder' Nearby events on the same curve: Codex · GPT-5.6 · SCPR · 15.49 h Not selected Best complete six-rung improvement at this time: +0.00% Not selected for a full ladder: both 5,000-update screens improved RPM (geometric gain 1.00074440 on E0; 1.00252280 on E2), but OEPR improved more. This is selection rejection, not a negative score. Trajectory steps: 2393, 2470, 2471 Codex · GPT-5.6 · MPR · 15.73 h Rejected / failed Best complete six-rung improvement at this time: +0.00% Rejected at both 5,000-update screens versus RPM: geometric gain 0.99291949 on E0 and 0.99619747 on E2. E0 also fails the macro safety gate. Trajectory steps: 2470, 2471, 2538, 2539, 2548Codex · GPT-5.6 · AGPR · 17.00 h Full ladder launched Best complete six-rung improvement at this time: +0.29% 16.9h passes E2 gate; 17.0h 'AGPR's six-rung ladder is submitted'Codex · GPT-5.6 · PSPR · 26.20 h Full ladder launched Best complete six-rung improvement at this time: +0.29% 26.1h E2 screen passes; 26.2h 'launching the six scored contracts as opt-014-pspr-full'Codex · GPT-5.6 · OEPR · 36.60 h Partial / projected result Best complete six-rung improvement at this time: +0.29% Partial / projected score: +0.35% (projected, never completed); the completed best is unchanged. 36.6h 'OEPR projects to reward 1.00354 with every safety gate passing—stronger than the staged RPM—but it is not stageable until its exact-final checkpoints and reload audit finish'; never mentioned after the 52.1h restartCodex · GPT-5.6 · CPSR · 55.20 h Full ladder launched Best complete six-rung improvement at this time: +0.60% 55.1h passes E2; 55.2h 'frozen CPSR full ladder is launched as opt-017-cpsr-full'Codex · GPT-5.6 · AGPR → control · 56.10 h Reduced control Best complete six-rung improvement at this time: +0.60% 56.1h AGPR E5 completes; 55.8h 'exact-35k PSPR-vs-AGPR gain 1.0003199' — AGPR is PSPR's reduced control; no six-rung reward ever statedCodex · GPT-5.6 · CPSR · 66.40 h Stopped Best complete six-rung improvement at this time: +0.60% 66.4h 'CPSR is now formally recorded as incomplete due to infrastructure and trusted launch cutoff'Codex · GPT-5.6 · AGST · 0.81 h Rejected / failed Best complete six-rung improvement at this time: +0.00% Rejected at E0, update 5,000: geometric gain 0.97835421 versus reduced AdamH (Paloma micro BPB 1.46582524 vs 1.43076432; fixed-window loss 4.07710351 vs 3.99814064). Its six-rung ladder was stopped; no complete-ladder score. Trajectory steps: 119, 120, 143Codex · GPT-5.6 · CA-TFCPR · 9.76 h Rejected / failed Best complete six-rung improvement at this time: +0.00% Rejected at E0, update 5,000: geometric gain 0.97556756 versus T-FCPR; macro BPB fails the 1% non-regression gate. E2 evaluation was interrupted, so it has no completed E2 screen. Trajectory steps: 1527, 1528, 1556, 1561Codex · GPT-5.6 · T-FCPR · 10.91 h Rejected / failed Best complete six-rung improvement at this time: +0.00% Rejected at E2, update 5,000: geometric gain 0.99955754 versus FCPR, despite passing the macro gate. E0 had narrowly passed at 1.00010275. Trajectory steps: 1527, 1679, 1680 Nearby events on the same curve: Codex · GPT-5.6 · TARF · 10.90 h Plan not launched Best complete six-rung improvement at this time: +0.00% 9.2h preregistered contingent on T-FCPR passing; T-FCPR rejected 10.9h → never implemented or launchedCodex · GPT-5.6 · TA-FCPR · 12.22 h Rejected / failed Best complete six-rung improvement at this time: +0.00% Rejected at both 5,000-update screens versus FCPR: geometric gain 0.98204696 on E0 and 0.98710723 on E2; both macro safety gates fail. Trajectory steps: 1852, 1856, 1912, 1915 Nearby events on the same curve: Codex · GPT-5.6 · FCPR · 12.30 h Paused Best complete six-rung improvement at this time: +0.00% 12.3h 'FCPR E0 25k comparison is negative … the pause is operational, not a substituted final verdict'; never resumed, never formally rejectedCodex · GPT-5.6 · SV-FCPR · 13.40 h Rejected / failed Best complete six-rung improvement at this time: +0.00% Rejected at both 5,000-update screens versus FCPR: geometric gain 0.99957022 on E0 and 0.99866946 on E2. Both macro safety gates pass, but neither screen improves the scored pair. Trajectory steps: 2005, 2105, 2122Codex · GPT-5.6 · GPRM · 14.32 h Rejected / failed Best complete six-rung improvement at this time: +0.00% Rejected at both 5,000-update screens versus RPM: geometric gain 0.99965743 on E0 and 0.99957609 on E2; both macro gates pass. Trajectory steps: 2214, 2218, 2282, 2283Codex · GPT-5.6 · SCPR · 15.49 h Not selected Best complete six-rung improvement at this time: +0.00% Not selected for a full ladder: both 5,000-update screens improved RPM (geometric gain 1.00074440 on E0; 1.00252280 on E2), but OEPR improved more. This is selection rejection, not a negative score. Trajectory steps: 2393, 2470, 2471 Nearby events on the same curve: Codex · GPT-5.6 · OEPR · 15.40 h Full ladder launched Best complete six-rung improvement at this time: +0.00% 15.4h 'OEPR is promoted to a frozen full ladder' Codex · GPT-5.6 · MPR · 15.73 h Rejected / failed Best complete six-rung improvement at this time: +0.00% Rejected at both 5,000-update screens versus RPM: geometric gain 0.99291949 on E0 and 0.99619747 on E2. E0 also fails the macro safety gate. Trajectory steps: 2470, 2471, 2538, 2539, 2548Codex · GPT-5.6 · MPR · 15.73 h Rejected / failed Best complete six-rung improvement at this time: +0.00% Rejected at both 5,000-update screens versus RPM: geometric gain 0.99291949 on E0 and 0.99619747 on E2. E0 also fails the macro safety gate. Trajectory steps: 2470, 2471, 2538, 2539, 2548 Nearby events on the same curve: Codex · GPT-5.6 · OEPR · 15.40 h Full ladder launched Best complete six-rung improvement at this time: +0.00% 15.4h 'OEPR is promoted to a frozen full ladder' Codex · GPT-5.6 · SCPR · 15.49 h Not selected Best complete six-rung improvement at this time: +0.00% Not selected for a full ladder: both 5,000-update screens improved RPM (geometric gain 1.00074440 on E0; 1.00252280 on E2), but OEPR improved more. This is selection rejection, not a negative score. Trajectory steps: 2393, 2470, 2471Codex · GPT-5.6 · CTPR · 34.90 h Rejected / failed Best complete six-rung improvement at this time: +0.29% Rejected at both 5,000-update screens versus PSPR. The later exact re-audit reports geometric gain 0.99970355 on E0 and 0.99900770 on E2. Earlier commentary used slightly different comparisons; these are the re-audited screen-to-screen values. Trajectory steps: 5258, 5326, 5881, 5885Codex · GPT-5.6 · PSGR · 53.51 h Rejected / failed Best complete six-rung improvement at this time: +0.60% Rejected by the two-screen promotion rule: at update 5,000 versus PSPR, E0 passes narrowly (1.00005941), but E2 is below one (0.99989665); macro safety passes. Trajectory steps: 6034, 6084, 6086Codex · GPT-5.6 · RPM · 16.18 h New measured best Best complete six-rung improvement at this time: +0.29% 16.2h 'RPM is now staged as the first eligible incumbent: exact-step reward 1.0029320791' Trajectory steps: 2612Codex · GPT-5.6 · PSPR · 52.19 h New measured best Best complete six-rung improvement at this time: +0.60% 52.2h 'PSPR's direct-step reward is 1.005988, a substantial improvement over the staged 1.002932 incumbent'; 56.1h 'PSPR is now the staged incumbent at reward 1.0059877526, replacing RPM' Trajectory steps: 5879RPM +0.29% — polar residual mixingPSPR +0.60% — principal-sine polar transportClaude Code · Opus 5 · RMBT v1 · 2.20 h Full ladder launched Best complete six-rung improvement at this time: +0.00% 2.2h 'Ladder launched from a single frozen source' Nearby events on the same curve: Claude Code · Opus 5 · Cautious gating · 1.77 h Rejected / failed Best complete six-rung improvement at this time: +0.00% E0 screen at 4,000 updates: fixed-window training loss 3.84154 versus 3.75773 for the mechanism-off control (+2.231%, worse). This is a single-rung screen, not a six-rung score. Trajectory steps: 142 Claude Code · Opus 5 · RMBT v1: cost overrun · 2.78 h Rejected / failed Best complete six-rung improvement at this time: +0.00% Initial ladder stopped before completion: E0/E1/E2 projected costs were 4.9%, 10.5%, and 9.1% above the AdamH manifest. E0 reached 5,000 updates and E5 reached zero; no complete six-rung score was measured. Trajectory steps: 196, 200, 201, 202Claude Code · Opus 5 · RMBT v2 · 3.00 h Full ladder launched Best complete six-rung improvement at this time: +0.00% 3.0h 'Ladder relaunched with the optimized source' (retraction rewritten as two Triton kernels) Nearby events on the same curve: Claude Code · Opus 5 · RMBT v1: cost overrun · 2.78 h Rejected / failed Best complete six-rung improvement at this time: +0.00% Initial ladder stopped before completion: E0/E1/E2 projected costs were 4.9%, 10.5%, and 9.1% above the AdamH manifest. E0 reached 5,000 updates and E5 reached zero; no complete six-rung score was measured. Trajectory steps: 196, 200, 201, 202Claude Code · Opus 5 · E5 relaunch · 6.80 h Infrastructure retry Best complete six-rung improvement at this time: +0.00% 6.7h 'E5 failed with zero updates' (NCCL mixed RoCE/IB); 6.8h relaunched as opt-002-rmbt-e5bClaude Code · Opus 5 · preview · 17.80 h Partial / projected result Best complete six-rung improvement at this time: +0.00% Partial / projected score: +0.29% (E5 still at 15k); the completed best is unchanged. 17.8h 'Six-rung preview reward: 1.00292' (E5 not yet at its 35,510 scoring update)Claude Code · Opus 5 · exhaustion · 59.23 h Stopped Best complete six-rung improvement at this time: +0.35% 59.2h 'Result — the optimizer works, but the ladder cannot be staged … No submission.json is producible — I am reporting exhaustion' Trajectory steps: 2297 Nearby events on the same curve: Claude Code · Opus 5 · RMBT final validation failed · 59.21 h Rejected / failed Best complete six-rung improvement at this time: +0.35% The handoff staging check again failed endpoint noninferiority and mechanism ablation on the same measured ladder. The session ended at 59.22549 hours (step 2297, 2026-09-03T05:06:48.820Z) reporting exhaustion and no submission.json. Trajectory steps: 2294, 2297Claude Code · Opus 5 · Cautious gating · 1.77 h Rejected / failed Best complete six-rung improvement at this time: +0.00% E0 screen at 4,000 updates: fixed-window training loss 3.84154 versus 3.75773 for the mechanism-off control (+2.231%, worse). This is a single-rung screen, not a six-rung score. Trajectory steps: 142 Nearby events on the same curve: Claude Code · Opus 5 · RMBT v1 · 2.20 h Full ladder launched Best complete six-rung improvement at this time: +0.00% 2.2h 'Ladder launched from a single frozen source'Claude Code · Opus 5 · RMBT v1: cost overrun · 2.78 h Rejected / failed Best complete six-rung improvement at this time: +0.00% Initial ladder stopped before completion: E0/E1/E2 projected costs were 4.9%, 10.5%, and 9.1% above the AdamH manifest. E0 reached 5,000 updates and E5 reached zero; no complete six-rung score was measured. Trajectory steps: 196, 200, 201, 202 Nearby events on the same curve: Claude Code · Opus 5 · RMBT v1 · 2.20 h Full ladder launched Best complete six-rung improvement at this time: +0.00% 2.2h 'Ladder launched from a single frozen source' Claude Code · Opus 5 · RMBT v2 · 3.00 h Full ladder launched Best complete six-rung improvement at this time: +0.00% 3.0h 'Ladder relaunched with the optimized source' (retraction rewritten as two Triton kernels)Claude Code · Opus 5 · Vocabulary redistribution · 3.93 h Rejected / failed Best complete six-rung improvement at this time: +0.00% E0 screen at 4,000 updates: applying trust redistribution to the vocabulary matrix raised fixed-window training loss to 3.82344 versus 3.75708 for the same-kernel mechanism-off control (+1.766%, worse). No six-rung score. Trajectory steps: 259 Nearby events on the same curve: Claude Code · Opus 5 · Overcorrection (alpha=2) · 4.05 h Rejected / failed Best complete six-rung improvement at this time: +0.00% E0 screen at 4,000 updates: alpha=2 raised fixed-window training loss to 3.77349 versus 3.75708 for the same-kernel mechanism-off control (+0.437%, worse). No six-rung score. Trajectory steps: 269Claude Code · Opus 5 · Overcorrection (alpha=2) · 4.05 h Rejected / failed Best complete six-rung improvement at this time: +0.00% E0 screen at 4,000 updates: alpha=2 raised fixed-window training loss to 3.77349 versus 3.75708 for the same-kernel mechanism-off control (+0.437%, worse). No six-rung score. Trajectory steps: 269 Nearby events on the same curve: Claude Code · Opus 5 · Vocabulary redistribution · 3.93 h Rejected / failed Best complete six-rung improvement at this time: +0.00% E0 screen at 4,000 updates: applying trust redistribution to the vocabulary matrix raised fixed-window training loss to 3.82344 versus 3.75708 for the same-kernel mechanism-off control (+1.766%, worse). No six-rung score. Trajectory steps: 259Claude Code · Opus 5 · RMBT staging failed · 41.11 h Rejected / failed Best complete six-rung improvement at this time: +0.35% The already measured RMBT ladder failed two staging guards: E1 cost-matched endpoint ratio 1.00614 exceeded the 1.005 limit, and E5 cost-matched mechanism-ablation gain 0.99908 fell below 1. This did not produce a new model score. Trajectory steps: 1589, 1590Claude Code · Opus 5 · RMBT staging retry failed · 53.51 h Rejected / failed Best complete six-rung improvement at this time: +0.35% After supplementary controls, the same frozen RMBT ladder still failed E1 endpoint noninferiority and mechanism ablation. The full candidate checkpoints and measured score did not change; no submission.json was produced. Trajectory steps: 2057, 2058, 2062Claude Code · Opus 5 · RMBT final validation failed · 59.21 h Rejected / failed Best complete six-rung improvement at this time: +0.35% The handoff staging check again failed endpoint noninferiority and mechanism ablation on the same measured ladder. The session ended at 59.22549 hours (step 2297, 2026-09-03T05:06:48.820Z) reporting exhaustion and no submission.json. Trajectory steps: 2294, 2297 Nearby events on the same curve: Claude Code · Opus 5 · exhaustion · 59.23 h Stopped Best complete six-rung improvement at this time: +0.35% 59.2h 'Result — the optimizer works, but the ladder cannot be staged … No submission.json is producible — I am reporting exhaustion' Trajectory steps: 2297Claude Code · Opus 5 · RMBT · 29.18 h New measured best Best complete six-rung improvement at this time: +0.35% 29.2h 'All six rungs are now at their exact scoring updates. Step-matched reward = 1.00355'; report.html: 'Exhaustion reported; no staging … two failed cost-based guards' Trajectory steps: 1170RMBT +0.35% — radius-matched block trustCodex · GPT-5.6Claude Code · Opus 5Marin Adam Baseline

Task

Design an optimizer that scales

Create a new optimizer mechanism and run it unchanged across six language models, from 550M to 2.545B parameters. Each rung is compared with a locked AdamH control on Paloma bits-per-byte and fixed-window training loss. A regression greater than 1% at any rung fails the submission.

Result. Codex (GPT-5.6) tested 17 hypotheses and produced PSPR, which lowered Paloma bits-per-byte at every rung: 1.0765 → 1.0681 at 550M and 0.9088 → 0.9018 at 2.545B, where fixed-window loss went 2.6040 → 2.5852. Per rung it is 0.09% to 0.93% ahead of the control. Claude Code (Opus 5) produced RMBT, which reached 0.9020 bits-per-byte and 2.5847 loss at 2.545B but regressed 0.21% at 837M. It met the scientific gates and missed two cost-matched staging guards.

Behaviour. Codex tested whether balancing matrix-update strength across directions could improve training, then rejected a stronger correction when it lost at both small-model screens. Claude rescaled each hidden unit's update relative to its weights; extending the rule to vocabulary rows worsened loss. When one model size contradicted the apparent benefit, Claude ran a fresh control with the mechanism disabled. Both tested at larger sizes, and Claude revised its claim that the benefit would grow with every increase in scale.

task ↗ Codex trajectory ↗ Claude trajectory ↗

Post-trainingData synthesize≈20,864 H100-hours / per run

Change the training data; keep the experiment fixed Both runs start from Qwen3.5-122B-A10B. The task package supplies 2,560 shared base tasks and 2,560 baseline tasks. The agent replaces the baseline tasks with 2,560 newly generated, verifiable RL tasks designed in 24 hours, 512 per capability. Both models train on 5,120 tasks with the same RL settings, then take the same five held-out benchmarks. Synthesized RL Data Post-trained Qwen3.5-122B-A10B Baseline Agent RL tasks Ref task samples Ref task samples + 2,560 human baseline tasks + 2,560 newly synthesized RL tasks Both pools from the task package 512 per capability Train with the same RL settings 5 held-out benchmarks
The task. The agent generates new, verifiable RL tasks in 24 hours to replace the baseline human task pool.
BenchmarkOriginalBaselineAgentΔ
PolyMath
Math
68.0668.4068.21+0.16
MMLU-Pro
MCQA
86.3086.2986.44+0.14
IFBench
Instruction following
67.0167.6967.35+0.34
LongBench V2
Long context
64.0265.2164.020.00
LiveCodeBench V6
Coding
72.0671.0973.54+1.48
Geometric mean of all five71.0971.3771.51+0.42
Release reference scores (%). Original is the starting post-trained model; Baseline and Agent use the same GRPO recipe. Δ is agent minus original, in percentage points, computed before rounding.

Task

Synthesize RL tasks to further optimize a 122B post-trained model

Generate and select 2,560 verifiable RL tasks in 24 hours to improve Qwen3.5-122B-A10B (already post-trained): 512 each for math, multiple-choice reasoning, instruction following, long context, and coding. Each task needs a prompt and an automatically checkable answer, constraint set, or test suite. The agent may change only the RL task data; the model, GRPO training recipe, and reward rules are fixed.

Result. The Claude Code agent's data reached 71.51 on the geometric mean of five benchmarks, versus 71.37 for the baseline. Compared with the baseline, coding improved most (+2.45 points); math, instruction following, and long context scored lower.

Evaluation. LiveCodeBench V6 averages 10 generated solutions per problem across 175 problems (1,750 responses) to reduce sampling variance. PolyMath (9,000 items), MMLU-Pro (12,032), IFBench (294), and LongBench V2 (503) use one response per item.

Behaviour. The Claude Code agent generated arithmetic questions with deliberately similar answer choices, layered writing constraints, long logs, and coding problems, with answers checked by calculations, rules, or tests. It repeatedly sampled the starting model and selected questions that produced both correct and incorrect answers, aiming to provide useful reward contrast. It adjusted noisy small-sample estimates and compared filters for empty or overly long responses. The strategy connects data selection to the training rule.

Why we designed it this way

01Why synthesized RL tasks only, and no SFT?

SFT data is easy to distill from Claude or GPT, which raises data-license problems. An RL task avoids this and it also lets the agent watch the policy model’s pass rate during training and set task difficulty to match.

02Why merge five domains?

One domain is easy to overfit. Math, multiple-choice reasoning, instruction following, long context and coding together are classic domains what a production post-training run has to balance, and the geometric mean over five benchmarks rewards gains that hold across all of them.

03Where does the baseline data come from?

All drawn from DAPO (math), Nemotron (multiple choice), multi-constraint instruction tasks, HotpotQA contexts (long context) and Open-R1 Codeforces problems (coding).

Vision-Gen5,815 H100-hours / per run

Fixed data. Open model. Train a text-to-image model from scratch on one pass over 100 million captioned GPIC images. The data and the single-pass budget are fixed; the architecture, objective and model size are open. The model generates 256 by 256 images for held-out captions at guidance 1, and a frozen FD-DINOv2 evaluator scores them. The score informs the next idea. The small pictures are schematic. Fixed data. Open model. 100M captioned images 1 pass Open model Design & train Generate 256 × 256 Fixed evaluator Guidance 1 FD-DINOv2 ↓ Iterate
100M images · one pass · open model design.Swipe to explore →
Codex: ten training launches, two completed passes Codex made ten training launches in ledger order. Four never read data: a launcher without resource directives, a cancelled retry, an invalid partition, and a text-encoder path failure. Four full-budget runs were killed at the 1,000 H100-hour cap short of a full pass, consuming 28.0, 68.0, 66.0 and 96.0 million images. A 45.8M-parameter model at batch 640 completed the 100 million-image pass and screened at FD-DINOv2 1114.03; restoring 142.8M parameters at the same batch completed too and screened at 778.61, 30.11% lower, and was selected. FD-DINOv2 is a local validation-reference screening score; lower is better. No final sealed-test score is supplied. Both screens use 50,000 frozen evaluation captions, guidance 1, and 50 Euler steps. The saved checkpoints contain 156,240 training steps (99,993,600 images), while both training runs report completing 100 million images. A full-budget attempt runs 6.6 to 7.8 hours on 128 H100s; generation and screening add about an hour. Hover over a launch for its full ledger details. Codex · 10 launches −30.11% FD Launch Images seen FD-DINOv2 ↓ 0 50M 100M Launch 001: launcher had no resource request. id: gpic-001-pure-conditional; cycle: 1; status: failed; gpu_hours: 0.0; images_seen: 0; line: 102; mechanism: Remove caption dropout on original 1.122B JiT; batch 256.; outcome: Launcher lacked required allocation directives; failed before training.; steps: None; parameters: 1122396928; global_batch: 256; consumed_training_data: False; fd_dinov2: None; full_pass_training_reported: False 01 No dropout Launch 002: cancelled before allocation. id: gpic-002-pure-conditional-retry; cycle: 3; status: failed; gpu_hours: 0.0; images_seen: 0; line: 114; mechanism: Same candidate; explicit 16-node allocation wrapper.; outcome: Cancelled before allocation; no training.; steps: None; parameters: 1122396928; global_batch: 256; consumed_training_data: False; fd_dinov2: None; full_pass_training_reported: False 02 New launcher Launch 003: invalid partition. id: gpic-003-pure-conditional-low; cycle: 5; status: failed; gpu_hours: 0.0; images_seen: 0; line: 126; mechanism: Same candidate; resubmit to low partition.; outcome: Invalid partition recovery cancelled before allocation; no training.; steps: None; parameters: 1122396928; global_batch: 256; consumed_training_data: False; fd_dinov2: None; full_pass_training_reported: False 03 New partition Launch 004: failed at 21 s, text encoder path. id: gpic-004-pure-conditional-seg01; cycle: 7; status: failed; gpu_hours: 0.7467; images_seen: 0; line: 138; mechanism: Same candidate; canonical P3 launcher.; outcome: Offline Qwen path unresolved; failed after 21s before reading data.; steps: None; parameters: 1122396928; global_batch: 256; consumed_training_data: False; fd_dinov2: None; full_pass_training_reported: False 04 Canonical launcher Launch 005: budget-killed · 3.9 steps/s. id: gpic-005-pure-conditional-seg01; cycle: 9; status: killed_budget; gpu_hours: 1000.96; images_seen: 27998208; line: 150; mechanism: Original 1.122B JiT with no caption dropout; fix offline text-encoder path.; outcome: Budget-killed; 109368 steps; no FD.; steps: 109368; parameters: 1122396928; global_batch: 256; consumed_training_data: True; fd_dinov2: None; full_pass_training_reported: False 05 Encoder fix 28M Launch 006: budget-killed · 2.4× throughput. id: gpic-006-compact-jit; cycle: 11; status: killed_budget; gpu_hours: 1000.1422; images_seen: 67995648; line: 162; mechanism: Shrink to 142.8M: hidden 768, 8 image blocks, 2 text blocks; retain batch256 and 300-token context.; outcome: Budget-killed; 2.43x step throughput still insufficient.; steps: 265608; parameters: 142846208; global_batch: 256; consumed_training_data: True; fd_dinov2: None; full_pass_training_reported: False 06 142.8M model 68M Launch 007: budget-killed · no gain. id: gpic-007-compact-shorttext; cycle: 13; status: killed_budget; gpu_hours: 1000.6756; images_seen: 65995776; line: 174; mechanism: Reduce caption context 300 to 64 tokens on compact JiT.; outcome: Budget-killed; throughput did not improve over gpic-006.; steps: 257796; parameters: 142664960; global_batch: 256; consumed_training_data: True; fd_dinov2: None; full_pass_training_reported: False 07 64-token captions 66M Launch 008: budget-killed · 4% short of a pass. id: gpic-008-compact-batch512; cycle: 15; status: killed_budget; gpu_hours: 1000.0711; images_seen: 95993856; line: 186; mechanism: Restore 300-token context; raise global batch256 to512.; outcome: Budget-killed; reached95.994M images, about4% short.; steps: 187488; parameters: 142846208; global_batch: 512; consumed_training_data: True; fd_dinov2: None; full_pass_training_reported: False 08 Batch 512 96M Launch 009: completed · screened. id: gpic-009-small-batch640; cycle: 17; status: completed; gpu_hours: 855.68; images_seen: 100000000; line: 198; mechanism: Shrink to45.8M denoiser; raise batch to640; target156250 steps.; outcome: Trainer completed100M images; stored checkpoint is156240 steps (99993600 images). FD1114.034971; superseded.; steps: 156250; parameters: 45835904; global_batch: 640; consumed_training_data: True; fd_dinov2: 1114.0349711301735; full_pass_training_reported: True; saved_checkpoint_steps: 156240; saved_checkpoint_images: 99993600 09 45.8M model 1114.03 Launch 010: completed · screened · selected. id: gpic-010-compact-batch640; cycle: 23; status: completed; gpu_hours: 839.9644; images_seen: 100000000; line: 234; mechanism: Restore142.8M capacity while keeping batch640 and one-pass schedule.; outcome: Trainer completed100M images; stored checkpoint156240; FD778.606203; selected.; steps: 156250; parameters: 142846208; global_batch: 640; consumed_training_data: True; fd_dinov2: 778.6062025358918; full_pass_training_reported: True; saved_checkpoint_steps: 156240; saved_checkpoint_images: 99993600 10 142.8M model 778.61 Stopped Full pass Selected
The research. Local screening; lower FD is better.Swipe to explore →

Task

Better generation from a single pass

Train a text-to-image model from scratch on 100M GPIC captioned images, with each example seen at most once per candidate. Architecture, objective and model size are open. Data exposure, 256×256 output, pure conditional sampling (guidance 1), and the FD-DINOv2 evaluator are fixed.

Codex ten launches. Four never read data: a launcher without resource directives, a cancelled retry, an invalid partition, and an unresolved text-encoder path. Four full-budget runs were killed at the 1,000 H100-hour cap short of a full pass, each one raising throughput, from 28M to 96M images. Shrinking to a 45.8M-parameter JiT at batch 640 completed the pass and screened at FD-DINOv2 1114.03; restoring 142.8M parameters at the same batch completed too, screened at 778.61 (−30.11%), and was selected.

Why we designed it this way

01Why isn’t a larger model automatically better?

Our small twist is to fix the data budget and leave the model open. Extra parameters only help if the recipe can learn effectively from the available examples. In NanoGPT Slowrun, a 1.4B model beat a 2.7B model without sufficient regularization; stronger regularization restored gains from scale. Its data-efficiency experiments also show model-size rankings reversing as the data budget changes. GPIC turns that idea into a generative-modeling question: which architecture, objective and training recipe make one pass count?

02Why only one epoch?

We want recipes for large-scale generative pretraining, where vast, diverse corpora make repeated full passes expensive. A strict single pass is a useful approximation to that operating regime: improvements must come from learning more per example. Training diffusion models for hundreds of epochs on ImageNet answers a different question and can favor recipes that do not carry over to this setting. The GPIC paper motivates precisely this shift toward large datasets and rich text conditioning. Here, every candidate gets the same 100M-image opportunity.

03Why fix guidance to 1?

Measure the model’s own conditional generation. Classifier-free guidance changes the sampling distribution, so tuning it can move FD without improving the trained model. We remove that extra tuning axis and keep room to study improvements in pure conditional quality. We also want an architecture-neutral task: diffusion, flow matching, autoregressive models and other generators should all compete without needing a CFG-specific recipe. Autoregressive models can use CFG too; the shared rule is that nobody uses guidance amplification. Guidance 1 still uses the caption—it is not unconditional generation.

Lightweight task samples

These samples are for an initial preview only and may be further adjusted. More samples are ongoing.

2 H100 · 6 hVision

Teaser figure from Molmo2
Molmo2 · source ↗
0.2750.3310.3860.4410.496Baseline · 0.3195 soft-F10246.5Elapsed time (h)Mean spatiotemporal soft-F1…Claude Fable 5 · submission 1 · 0.32 · 0.0951 h, Judge result recordedClaude Fable 5 · submission 2 · 0.358 · 0.195 h, Judge result recordedClaude Fable 5 · submission 3 · 0.334 · 0.29 h, Judge result recordedClaude Fable 5 · submission 4 · 0.358 · 0.387 h, Judge result recordedClaude Fable 5 · submission 5 · 0.358 · 0.48 h, Judge result recordedClaude Fable 5 · submission 6 · 0.338 · 0.576 h, Judge result recordedClaude Fable 5 · submission 7 · 0.358 · 0.67 h, Judge result recordedClaude Fable 5 · submission 8 · 0.358 · 0.764 h, Judge result recordedClaude Fable 5 · submission 9 · 0.348 · 0.851 h, Judge result recordedClaude Fable 5 · submission 10 · 0.435 · 0.946 h, Judge result recordedClaude Fable 5 · submission 11 · 0.426 · 1.04 h, Judge result recordedClaude Fable 5 · submission 12 · 0.359 · 1.13 h, Judge result recordedClaude Fable 5 · submission 13 · 0.35 · 1.22 h, Judge result recordedClaude Fable 5 · submission 14 · 0.427 · 1.31 h, Judge result recordedClaude Fable 5 · submission 15 · 0.378 · 1.39 h, Judge result recordedClaude Fable 5 · submission 16 · 0.378 · 1.49 h, Judge result recordedClaude Fable 5 · submission 17 · 0.431 · 1.57 h, Judge result recordedClaude Fable 5 · submission 18 · 0.444 · 1.66 h, Judge result recordedClaude Fable 5 · submission 19 · 0.41 · 1.76 h, Judge result recordedClaude Fable 5 · submission 20 · 0.367 · 1.85 h, Judge result recordedClaude Fable 5 · submission 21 · 0.43 · 1.94 h, Judge result recordedClaude Fable 5 · submission 22 · 0.438 · 2.02 h, Judge result recordedClaude Fable 5 · submission 23 · 0.444 · 2.11 h, Judge result recordedClaude Fable 5 · submission 24 · 0.444 · 2.19 h, Judge result recordedClaude Fable 5 · submission 25 · 0.444 · 2.28 h, Judge result recordedClaude Fable 5 · submission 26 · 0.427 · 2.37 h, Judge result recordedClaude Fable 5 · submission 27 · 0.309 · 2.45 h, Judge result recordedClaude Fable 5 · submission 28 · 0.309 · 2.53 h, Judge result recordedClaude Fable 5 · submission 29 · 0.411 · 2.61 h, Judge result recordedClaude Fable 5 · submission 30 · 0.444 · 2.7 h, Judge result recordedClaude Fable 5 · submission 31 · 0.378 · 2.79 h, Judge result recordedClaude Fable 5 · submission 32 · 0.384 · 2.87 h, Judge result recordedClaude Fable 5 · submission 33 · 0.424 · 2.95 h, Judge result recordedClaude Fable 5 · submission 34 · 0.403 · 3.03 h, Judge result recordedClaude Fable 5 · submission 35 · 0.444 · 3.12 h, Judge result recordedClaude Fable 5 · submission 36 · 0.431 · 3.26 h, Judge result recordedClaude Fable 5 · submission 37 · 0.427 · 3.35 h, Judge result recordedClaude Fable 5 · submission 38 · 0.426 · 3.43 h, Judge result recordedClaude Fable 5 · submission 39 · 0.444 · 3.51 h, Judge result recordedClaude Fable 5 · submission 40 · 0.349 · 3.61 h, Judge result recordedClaude Fable 5 · submission 41 · 0.342 · 3.7 h, Judge result recordedClaude Fable 5 · submission 42 · 0.448 · 3.78 h, Judge result recordedClaude Fable 5 · submission 43 · 0.431 · 3.86 h, Judge result recordedClaude Fable 5 · submission 44 · 0.402 · 3.94 h, Judge result recordedClaude Fable 5 · submission 45 · 0.444 · 4.02 h, Judge result recordedClaude Fable 5 · submission 46 · 0.392 · 4.1 h, Judge result recordedClaude Fable 5 · submission 47 · 0.431 · 4.18 h, Judge result recordedClaude Fable 5 · submission 48 · 0.419 · 4.26 h, Judge result recordedClaude Fable 5 · submission 49 · 0.462 · 4.34 h, Judge result recordedClaude Fable 5 · submission 50 · 0.448 · 4.41 h, Judge result recordedClaude Fable 5 · submission 51 · 0.434 · 4.5 h, Judge result recordedClaude Fable 5 · submission 52 · 0.434 · 4.57 h, Judge result recordedClaude Fable 5 · submission 53 · 0.419 · 4.65 h, Judge result recordedClaude Fable 5 · submission 54 · 0.414 · 4.73 h, Judge result recordedClaude Fable 5 · submission 55 · 0.406 · 4.81 h, Judge result recordedClaude Fable 5 · submission 56 · 0.462 · 4.89 h, Judge result recordedClaude Fable 5 · submission 57 · 0.376 · 4.98 h, Judge result recordedClaude Fable 5 · submission 58 · 0.405 · 5.06 h, Judge result recordedClaude Fable 5 · submission 59 · 0.362 · 5.14 h, Judge result recordedClaude Fable 5 · 0.462GPT-5.6 Sol · submission 2 · 0.334 · 0.553 h, Judge result recordedGPT-5.6 Sol · submission 3 · 0.32 · 0.651 h, Judge result recordedGPT-5.6 Sol · submission 4 · 0.352 · 0.783 h, Judge result recordedGPT-5.6 Sol · submission 5 · 0.378 · 0.874 h, Judge result recordedGPT-5.6 Sol · submission 6 · 0.42 · 0.974 h, Judge result recordedGPT-5.6 Sol · submission 7 · 0.434 · 1.08 h, Judge result recordedGPT-5.6 Sol · submission 8 · 0.341 · 1.17 h, Judge result recordedGPT-5.6 Sol · submission 9 · 0.434 · 1.31 h, Judge result recordedGPT-5.6 Sol · submission 10 · 0.414 · 1.41 h, Judge result recordedGPT-5.6 Sol · submission 11 · 0.431 · 1.52 h, Judge result recordedGPT-5.6 Sol · submission 12 · 0.426 · 1.69 h, Judge result recordedGPT-5.6 Sol · submission 13 · 0.391 · 1.87 h, Judge result recordedGPT-5.6 Sol · submission 14 · 0.405 · 1.96 h, Judge result recordedGPT-5.6 Sol · submission 15 · 0.418 · 2.06 h, Judge result recordedGPT-5.6 Sol · submission 16 · 0.462 · 2.15 h, Judge result recordedGPT-5.6 Sol · submission 17 · 0.406 · 2.24 h, Judge result recordedGPT-5.6 Sol · submission 18 · 0.448 · 2.36 h, Judge result recordedGPT-5.6 Sol · submission 19 · 0.419 · 2.49 h, Judge result recordedGPT-5.6 Sol · submission 20 · 0.419 · 2.59 h, Judge result recordedGPT-5.6 Sol · submission 21 · 0.361 · 2.77 h, Judge result recordedGPT-5.6 Sol · submission 22 · 0.406 · 2.87 h, Judge result recordedGPT-5.6 Sol · submission 23 · 0.369 · 3 h, Judge result recordedGPT-5.6 Sol · submission 24 · 0.367 · 3.1 h, Judge result recordedGPT-5.6 Sol · submission 25 · 0.438 · 3.3 h, Judge result recordedGPT-5.6 Sol · submission 26 · 0.411 · 3.39 h, Judge result recordedGPT-5.6 Sol · submission 27 · 0.431 · 3.49 h, Judge result recordedGPT-5.6 Sol · submission 28 · 0.376 · 3.62 h, Judge result recordedGPT-5.6 Sol · submission 29 · 0.424 · 3.71 h, Judge result recordedGPT-5.6 Sol · submission 30 · 0.426 · 3.81 h, Judge result recordedGPT-5.6 Sol · submission 31 · 0.392 · 3.9 h, Judge result recordedGPT-5.6 Sol · submission 32 · 0.341 · 4 h, Judge result recordedGPT-5.6 Sol · submission 33 · 0.403 · 4.1 h, Judge result recordedGPT-5.6 Sol · submission 34 · 0.444 · 4.24 h, Judge result recordedGPT-5.6 Sol · submission 35 · 0.349 · 4.34 h, Judge result recordedGPT-5.6 Sol · submission 36 · 0.43 · 4.48 h, Judge result recordedGPT-5.6 Sol · submission 37 · 0.427 · 4.58 h, Judge result recordedGPT-5.6 Sol · submission 38 · 0.384 · 4.68 h, Judge result recordedGPT-5.6 Sol · submission 39 · 0.435 · 4.78 h, Judge result recordedGPT-5.6 Sol · submission 40 · 0.41 · 4.88 h, Judge result recordedGPT-5.6 Sol · submission 41 · 0.367 · 4.99 h, Judge result recordedGPT-5.6 Sol · submission 42 · 0.402 · 5.09 h, Judge result recordedGPT-5.6 Sol · submission 43 · 0.403 · 5.18 h, Judge result recordedGPT-5.6 Sol · submission 44 · 0.391 · 5.28 h, Judge result recordedGPT-5.6 Sol · submission 45 · 0.358 · 5.37 h, Judge result recordedGPT-5.6 Sol · submission 46 · 0.356 · 5.46 h, Judge result recordedGPT-5.6 Sol · submission 47 · 0.408 · 5.63 h, Judge result recordedGPT-5.6 Sol · submission 48 · 0.333 · 5.72 h, Judge result recordedGPT-5.6 Sol · submission 49 · 0.327 · 5.81 h, Judge result recordedGPT-5.6 Sol · submission 50 · 0.309 · 5.92 h, Judge result recordedGPT-5.6 Sol · submission 51 · 0.363 · 6.02 h, Judge result recordedGPT-5.6 Sol · submission 52 · 0.339 · 6.11 h, Judge result recordedGPT-5.6 Sol · submission 53 · 0.361 · 6.21 h, Judge result recordedGPT-5.6 Sol · submission 54 · 0.337 · 6.3 h, Judge result recordedGPT-5.6 Sol · submission 55 · 0.406 · 6.39 h, Judge result recordedGPT-5.6 Sol · 0.462
submissionrunning bestbaseline · Judge measuredTime since run start · points mark recorded Judge results

Molmo2 Video-Pointing Inference Strategy

Improve how a frozen Molmo2 model points to anomalies in video by tuning which frames it sees, how it is prompted, and how long it can answer.

Judged on mean spatiotemporal pointing soft-F1, scored absolutely.

full task page → task ↗ run log ↗

8 H100 · 8 hPosttrain

Teaser figure from Small-Model Learnability Gap
Small-Model Learnability Gap · source ↗
3438424650Baseline · 46.073 pp02467.9Elapsed time (h)5-task macro accuracyClaude Fable 5 · submission 1 · 46.1 · 0.29 h, Judge result recordedClaude Fable 5 · submission 2 · 0 · 0.789 h, Judge result recordedClaude Fable 5 · submission 3 · 0 · 1.71 h, Judge result recordedClaude Fable 5 · submission 4 · 44.2 · 2.19 h, Judge result recordedClaude Fable 5 · submission 5 · 0 · 2.42 h, Judge result recordedClaude Fable 5 · submission 6 · 43 · 2.7 h, Judge result recordedClaude Fable 5 · submission 7 · 36.4 · 3.1 h, Judge result recordedClaude Fable 5 · submission 8 · 44.2 · 3.38 h, Judge result recordedClaude Fable 5 · submission 9 · 44.2 · 3.67 h, Judge result recordedClaude Fable 5 · submission 10 · 45.5 · 4.01 h, Judge result recordedClaude Fable 5 · submission 11 · 43.9 · 4.28 h, Judge result recordedClaude Fable 5 · submission 12 · 45.4 · 4.62 h, Judge result recordedClaude Fable 5 · submission 13 · 41.1 · 4.93 h, Judge result recordedClaude Fable 5 · submission 14 · 44.7 · 5.19 h, Judge result recordedClaude Fable 5 · submission 15 · 43.3 · 5.45 h, Judge result recordedClaude Fable 5 · submission 16 · 42.3 · 5.72 h, Judge result recordedClaude Fable 5 · submission 17 · 42.2 · 5.99 h, Judge result recordedClaude Fable 5 · submission 18 · 44.6 · 6.24 h, Judge result recordedClaude Fable 5 · submission 19 · 44 · 6.49 h, Judge result recordedClaude Fable 5 · submission 20 · 47.5 · 6.75 h, Judge result recordedClaude Fable 5 · submission 21 · 42.6 · 7.15 h, Judge result recordedClaude Fable 5 · submission 22 · 44 · 7.41 h, Judge result recordedClaude Fable 5 · submission 23 · 45.7 · 7.67 h, Judge result recordedClaude Fable 5 · 47.5GPT-5.6 Sol · submission 1 · 0 · 0.22 h, Judge result recordedGPT-5.6 Sol · submission 2 · 0 · 0.239 h, Judge result recordedGPT-5.6 Sol · submission 3 · 0 · 0.246 h, Judge result recordedGPT-5.6 Sol · submission 4 · 0 · 0.433 h, Judge result recordedGPT-5.6 Sol · submission 5 · 45.5 · 0.645 h, Judge result recordedGPT-5.6 Sol · submission 6 · 44.6 · 0.971 h, Judge result recordedGPT-5.6 Sol · submission 7 · 45.1 · 1.31 h, Judge result recordedGPT-5.6 Sol · submission 8 · 42 · 1.63 h, Judge result recordedGPT-5.6 Sol · submission 9 · 43.2 · 1.95 h, Judge result recordedGPT-5.6 Sol · submission 10 · 44.7 · 2.25 h, Judge result recordedGPT-5.6 Sol · submission 11 · 43.5 · 2.56 h, Judge result recordedGPT-5.6 Sol · submission 12 · 41.9 · 2.87 h, Judge result recordedGPT-5.6 Sol · submission 13 · 44.2 · 3.18 h, Judge result recordedGPT-5.6 Sol · submission 14 · 45 · 3.48 h, Judge result recordedGPT-5.6 Sol · submission 15 · 46.1 · 3.8 h, Judge result recordedGPT-5.6 Sol · submission 16 · 45.7 · 4.13 h, Judge result recordedGPT-5.6 Sol · submission 17 · 44.1 · 4.44 h, Judge result recordedGPT-5.6 Sol · submission 18 · 43.9 · 4.77 h, Judge result recordedGPT-5.6 Sol · submission 19 · 43.8 · 5.09 h, Judge result recordedGPT-5.6 Sol · submission 20 · 45.1 · 5.46 h, Judge result recordedGPT-5.6 Sol · submission 21 · 43.7 · 5.78 h, Judge result recordedGPT-5.6 Sol · submission 22 · 44.8 · 6.12 h, Judge result recordedGPT-5.6 Sol · submission 23 · 43.3 · 6.49 h, Judge result recordedGPT-5.6 Sol · submission 24 · 43.7 · 6.83 h, Judge result recordedGPT-5.6 Sol · submission 25 · 45.1 · 7.14 h, Judge result recordedGPT-5.6 Sol · submission 26 · 44.3 · 7.45 h, Judge result recordedGPT-5.6 Sol · submission 27 · 41.6 · 7.78 h, Judge result recordedGPT-5.6 Sol · 46.1
submissionrejected submissionrunning bestbaseline · Judge measuredTime since run start · points mark recorded Judge results

Learnability-Aware Long/Short CoT Adaptation

Adapt the supplied long and short mathematical reasoning examples so the same small language model learns more effectively under a fixed training recipe.

Judged on five-benchmark mathematical-reasoning accuracy, scored absolutely.

full task page → task ↗ run log ↗

2 H100 · 6 hSystems

Teaser figure from MInference 1.0
MInference 1.0 · source ↗
0.8811.081.271.461.66Baseline · 1.000×0246.4Elapsed time (h)18-case geometric-mean speedupGPT-5.6 Sol · submission 1 · 1.08 · 0.167 h, Judge result recordedGPT-5.6 Sol · submission 2 · 0 · 0.267 h, Judge result recordedGPT-5.6 Sol · submission 3 · 1.14 · 0.318 h, Judge result recordedGPT-5.6 Sol · submission 4 · 1.14 · 0.385 h, Judge result recordedGPT-5.6 Sol · submission 5 · 1.19 · 0.455 h, Judge result recordedGPT-5.6 Sol · submission 6 · 1.2 · 0.569 h, Judge result recordedGPT-5.6 Sol · submission 7 · 1.24 · 0.884 h, Judge result recordedGPT-5.6 Sol · submission 8 · 1.31 · 0.935 h, Judge result recordedGPT-5.6 Sol · submission 9 · 1.33 · 1.15 h, Judge result recordedGPT-5.6 Sol · submission 10 · 1.33 · 1.25 h, Judge result recordedGPT-5.6 Sol · submission 11 · 1.34 · 1.34 h, Judge result recordedGPT-5.6 Sol · submission 12 · 1.36 · 1.41 h, Judge result recordedGPT-5.6 Sol · submission 13 · 1.35 · 1.52 h, Judge result recordedGPT-5.6 Sol · submission 14 · 1.36 · 1.59 h, Judge result recordedGPT-5.6 Sol · submission 15 · 1.36 · 1.66 h, Judge result recordedGPT-5.6 Sol · submission 16 · 1.36 · 1.71 h, Judge result recordedGPT-5.6 Sol · submission 17 · 1.39 · 1.79 h, Judge result recordedGPT-5.6 Sol · submission 18 · 1.39 · 1.9 h, Judge result recordedGPT-5.6 Sol · submission 19 · 1.39 · 1.99 h, Judge result recordedGPT-5.6 Sol · submission 20 · 1.4 · 2.06 h, Judge result recordedGPT-5.6 Sol · submission 21 · 1.39 · 2.11 h, Judge result recordedGPT-5.6 Sol · submission 22 · 1.38 · 2.17 h, Judge result recordedGPT-5.6 Sol · submission 23 · 1.46 · 2.23 h, Judge result recordedGPT-5.6 Sol · submission 24 · 1.46 · 2.28 h, Judge result recordedGPT-5.6 Sol · submission 25 · 1.46 · 2.38 h, Judge result recordedGPT-5.6 Sol · submission 26 · 1.47 · 2.5 h, Judge result recordedGPT-5.6 Sol · submission 27 · 1.45 · 2.54 h, Judge result recordedGPT-5.6 Sol · submission 28 · 1.47 · 2.62 h, Judge result recordedGPT-5.6 Sol · submission 29 · 1.46 · 2.68 h, Judge result recordedGPT-5.6 Sol · submission 30 · 0 · 3.04 h, Judge result recordedGPT-5.6 Sol · submission 31 · 1.47 · 3.08 h, Judge result recordedGPT-5.6 Sol · submission 32 · 1.46 · 3.17 h, Judge result recordedGPT-5.6 Sol · submission 33 · 1.46 · 3.22 h, Judge result recordedGPT-5.6 Sol · submission 34 · 1.46 · 3.3 h, Judge result recordedGPT-5.6 Sol · submission 35 · 1.46 · 3.4 h, Judge result recordedGPT-5.6 Sol · submission 36 · 1.47 · 3.47 h, Judge result recordedGPT-5.6 Sol · submission 37 · 1.47 · 3.56 h, Judge result recordedGPT-5.6 Sol · submission 38 · 1.47 · 3.61 h, Judge result recordedGPT-5.6 Sol · submission 39 · 1.48 · 3.72 h, Judge result recordedGPT-5.6 Sol · submission 40 · 1.47 · 3.82 h, Judge result recordedGPT-5.6 Sol · submission 41 · 1.48 · 3.87 h, Judge result recordedGPT-5.6 Sol · submission 42 · 1.47 · 3.93 h, Judge result recordedGPT-5.6 Sol · submission 43 · 1.48 · 4 h, Judge result recordedGPT-5.6 Sol · submission 44 · 1.48 · 4.04 h, Judge result recordedGPT-5.6 Sol · submission 45 · 1.47 · 4.1 h, Judge result recordedGPT-5.6 Sol · submission 46 · 1.48 · 4.18 h, Judge result recordedGPT-5.6 Sol · submission 47 · 1.48 · 4.23 h, Judge result recordedGPT-5.6 Sol · submission 48 · 1.48 · 4.29 h, Judge result recordedGPT-5.6 Sol · submission 49 · 1.47 · 4.36 h, Judge result recordedGPT-5.6 Sol · submission 50 · 1.47 · 4.43 h, Judge result recordedGPT-5.6 Sol · submission 51 · 1.47 · 4.72 h, Judge result recordedGPT-5.6 Sol · submission 52 · 1.53 · 4.91 h, Judge result recordedGPT-5.6 Sol · submission 53 · 1.54 · 5.04 h, Judge result recordedGPT-5.6 Sol · submission 54 · 1.54 · 5.16 h, Judge result recordedGPT-5.6 Sol · submission 55 · 1.53 · 5.54 h, Judge result recordedGPT-5.6 Sol · submission 56 · 1.53 · 5.61 h, Judge result recordedGPT-5.6 Sol · submission 57 · 1.54 · 5.69 h, Judge result recordedGPT-5.6 Sol · submission 58 · 1.53 · 5.8 h, Judge result recordedGPT-5.6 Sol · submission 59 · 1.54 · 5.95 h, Judge result recordedGPT-5.6 Sol · submission 60 · 1.54 · 6.26 h, Judge result recordedGPT-5.6 Sol · submission 61 · 1.54 · 6.31 h, Judge result recordedGPT-5.6 Sol · 1.54Claude Opus 5 · submission 1 · 1.23 · 0.483 h, Judge result recordedClaude Opus 5 · submission 2 · 0 · 1.12 h, Judge result recordedClaude Opus 5 · submission 3 · 1.37 · 1.22 h, Judge result recordedClaude Opus 5 · submission 4 · 1.45 · 1.34 h, Judge result recordedClaude Opus 5 · submission 5 · 1.47 · 3.58 h, Judge result recordedClaude Opus 5 · submission 6 · 1.46 · 3.77 h, Judge result recordedClaude Opus 5 · submission 7 · 1.47 · 3.92 h, Judge result recordedClaude Opus 5 · submission 8 · 1.5 · 4.1 h, Judge result recordedClaude Opus 5 · submission 9 · 1.49 · 4.25 h, Judge result recordedClaude Opus 5 · submission 10 · 1.49 · 4.31 h, Judge result recordedClaude Opus 5 · submission 11 · 1.49 · 4.54 h, Judge result recordedClaude Opus 5 · submission 12 · 1.5 · 4.7 h, Judge result recordedClaude Opus 5 · submission 13 · 1.5 · 4.94 h, Judge result recordedClaude Opus 5 · submission 14 · 1.5 · 5.56 h, Judge result recordedClaude Opus 5 · submission 15 · 1.49 · 6.02 h, Judge result recordedClaude Opus 5 · 1.5
submissionrejected submissionrunning bestbaseline · paired normalizedTime since run start · points mark recorded Judge results

MInference 32-Head Sparse Prefill Kernel

Speed up a fixed sparse-attention prefill operator without changing which tokens it attends to or weakening its numerical result.

Judged on paired sparse-prefill speedup, measured against the reference in the same Judge run.

full task page → task ↗ run log ↗

2 H100 · 24 hAgents

Teaser figure from Agentic Context Engineering
Agentic Context Engineering · source ↗
0.5290.5880.6470.7070.766Baseline · 0.570 accuracy081624.1Elapsed time (h)Formula exact-answer accuracyClaude Opus 5 · submission 1 · 0.565 · 0.138 h, Judge result recordedClaude Opus 5 · submission 3 · 0.61 · 3.14 h, Judge result recordedClaude Opus 5 · submission 4 · 0.61 · 3.43 h, Judge result recordedClaude Opus 5 · submission 5 · 0.72 · 4.56 h, Judge result recordedClaude Opus 5 · submission 6 · 0.72 · 5.12 h, Judge result recordedClaude Opus 5 · submission 7 · 0.72 · 5.34 h, Judge result recordedClaude Opus 5 · submission 8 · 0.72 · 5.83 h, Judge result recordedClaude Opus 5 · submission 9 · 0.72 · 7.05 h, Judge result recordedClaude Opus 5 · submission 10 · 0.72 · 8.07 h, Judge result recordedClaude Opus 5 · submission 11 · 0.72 · 8.5 h, Judge result recordedClaude Opus 5 · submission 12 · 0.72 · 9.03 h, Judge result recordedClaude Opus 5 · submission 13 · 0.72 · 10.2 h, Judge result recordedClaude Opus 5 · submission 14 · 0.73 · 11.4 h, Judge result recordedClaude Opus 5 · submission 15 · 0.72 · 11.5 h, Judge result recordedClaude Opus 5 · submission 16 · 0.72 · 12.5 h, Judge result recordedClaude Opus 5 · submission 17 · 0.72 · 13.4 h, Judge result recordedClaude Opus 5 · submission 18 · 0.72 · 14.4 h, Judge result recordedClaude Opus 5 · submission 19 · 0.72 · 15.3 h, Judge result recordedClaude Opus 5 · submission 20 · 0.72 · 15.7 h, Judge result recordedClaude Opus 5 · submission 21 · 0.73 · 17.3 h, Judge result recordedClaude Opus 5 · submission 22 · 0.73 · 18 h, Judge result recordedClaude Opus 5 · submission 23 · 0.73 · 18.2 h, Judge result recordedClaude Opus 5 · submission 24 · 0.73 · 19 h, Judge result recordedClaude Opus 5 · submission 25 · 0.73 · 20.1 h, Judge result recordedClaude Opus 5 · submission 26 · 0.73 · 20.3 h, Judge result recordedClaude Opus 5 · submission 27 · 0.73 · 20.6 h, Judge result recordedClaude Opus 5 · submission 28 · 0.73 · 21.6 h, Judge result recordedClaude Opus 5 · submission 29 · 0.73 · 21.8 h, Judge result recordedClaude Opus 5 · submission 30 · 0.73 · 23 h, Judge result recordedClaude Opus 5 · 0.73GPT-5.6 Sol · submission 1 · 0.57 · 1.46 h, Judge result recordedGPT-5.6 Sol · submission 3 · 0.6 · 2.33 h, Judge result recordedGPT-5.6 Sol · submission 4 · 0.65 · 3.59 h, Judge result recordedGPT-5.6 Sol · submission 5 · 0.665 · 5.46 h, Judge result recordedGPT-5.6 Sol · submission 6 · 0.685 · 7.58 h, Judge result recordedGPT-5.6 Sol · submission 7 · 0.685 · 8.81 h, Judge result recordedGPT-5.6 Sol · submission 8 · 0.68 · 10 h, Judge result recordedGPT-5.6 Sol · submission 9 · 0.665 · 12.8 h, Judge result recordedGPT-5.6 Sol · submission 10 · 0.705 · 13.7 h, Judge result recordedGPT-5.6 Sol · submission 11 · 0.685 · 15.1 h, Judge result recordedGPT-5.6 Sol · submission 12 · 0.695 · 16.3 h, Judge result recordedGPT-5.6 Sol · submission 13 · 0.68 · 17.5 h, Judge result recordedGPT-5.6 Sol · submission 14 · 0.7 · 18.5 h, Judge result recordedGPT-5.6 Sol · submission 15 · 0.68 · 19.9 h, Judge result recordedGPT-5.6 Sol · submission 16 · 0.67 · 20.8 h, Judge result recordedGPT-5.6 Sol · submission 17 · 0.67 · 21.9 h, Judge result recordedGPT-5.6 Sol · submission 18 · 0.69 · 23 h, Judge result recordedGPT-5.6 Sol · 0.705
submissionrunning bestbaseline · Judge measuredTime since run start · points mark recorded Judge results

ACE Playbook Inspection and Repair

Inspect and repair the reusable advice in an ACE playbook so a frozen Qwen model answers Formula reasoning problems more accurately.

Judged on held-out Formula exact-answer accuracy, scored absolutely.

full task page → task ↗ run log ↗

3 H100 · 24 hVision

Teaser figure from MolmoWeb
MolmoWeb · source ↗
0.5920.5960.60.6040.608Baseline · 0.5966 action score081624.3Elapsed time (h)Macro action scoreClaude Opus 5 · submission 1 · 0.5966 · 0.784 h, Judge result recordedClaude Opus 5 · submission 2 · 0.5983 · 1.41 h, Judge result recordedClaude Opus 5 · submission 5 · 0.5932 · 3.17 h, Judge result recordedClaude Opus 5 · submission 7 · 0.5965 · 4.4 h, Judge result recordedClaude Opus 5 · submission 8 · 0.6057 · 5 h, Judge result recordedClaude Opus 5 · submission 9 · 0.5997 · 5.6 h, Judge result recordedClaude Opus 5 · submission 10 · 0.5934 · 6.21 h, Judge result recordedClaude Opus 5 · submission 11 · 0.6036 · 6.81 h, Judge result recordedClaude Opus 5 · submission 12 · 0.6040 · 7.41 h, Judge result recordedClaude Opus 5 · submission 13 · 0.6033 · 8.03 h, Judge result recordedClaude Opus 5 · submission 14 · 0.5973 · 8.63 h, Judge result recordedClaude Opus 5 · submission 15 · 0.6057 · 9.23 h, Judge result recordedClaude Opus 5 · submission 16 · 0.6013 · 9.87 h, Judge result recordedClaude Opus 5 · submission 19 · 0.6057 · 11.7 h, Judge result recordedClaude Opus 5 · submission 20 · 0.6057 · 12.3 h, Judge result recordedClaude Opus 5 · submission 23 · 0.6057 · 14.1 h, Judge result recordedClaude Opus 5 · submission 26 · 0.6057 · 16 h, Judge result recordedClaude Opus 5 · submission 27 · 0.5959 · 16.6 h, Judge result recordedClaude Opus 5 · submission 28 · 0.6057 · 17.3 h, Judge result recordedClaude Opus 5 · submission 29 · 0.5993 · 18 h, Judge result recordedClaude Opus 5 · submission 30 · 0.6057 · 18.5 h, Judge result recordedClaude Opus 5 · submission 32 · 0.6057 · 19.8 h, Judge result recordedClaude Opus 5 · submission 34 · 0.6057 · 21 h, Judge result recordedClaude Opus 5 · submission 35 · 0.6044 · 21.6 h, Judge result recordedClaude Opus 5 · submission 37 · 0.6057 · 22.8 h, Judge result recordedClaude Opus 5 · submission 39 · 0.6057 · 24.1 h, Judge result recordedClaude Opus 5 · 0.6057GPT-5.6 Sol · submission 1 · 0.5966 · 0.58 h, Judge result recordedGPT-5.6 Sol · submission 2 · 0.5985 · 1.78 h, Judge result recordedGPT-5.6 Sol · submission 4 · 0.5940 · 3.01 h, Judge result recordedGPT-5.6 Sol · submission 5 · 0.5985 · 3.63 h, Judge result recordedGPT-5.6 Sol · submission 6 · 0.5975 · 4.22 h, Judge result recordedGPT-5.6 Sol · submission 7 · 0.5991 · 4.81 h, Judge result recordedGPT-5.6 Sol · submission 8 · 0.6011 · 5.4 h, Judge result recordedGPT-5.6 Sol · submission 9 · 0.5959 · 6.01 h, Judge result recordedGPT-5.6 Sol · submission 10 · 0.6001 · 6.68 h, Judge result recordedGPT-5.6 Sol · submission 11 · 0.6011 · 7.27 h, Judge result recordedGPT-5.6 Sol · submission 13 · 0.6009 · 8.45 h, Judge result recordedGPT-5.6 Sol · submission 14 · 0.6011 · 9.04 h, Judge result recordedGPT-5.6 Sol · submission 15 · 0.6020 · 9.65 h, Judge result recordedGPT-5.6 Sol · submission 16 · 0.6022 · 10.2 h, Judge result recordedGPT-5.6 Sol · submission 20 · 0.6020 · 12.3 h, Judge result recordedGPT-5.6 Sol · submission 21 · 0.6031 · 12.9 h, Judge result recordedGPT-5.6 Sol · submission 25 · 0.6044 · 15.3 h, Judge result recordedGPT-5.6 Sol · submission 26 · 0.6028 · 15.9 h, Judge result recordedGPT-5.6 Sol · submission 27 · 0.6044 · 16.5 h, Judge result recordedGPT-5.6 Sol · submission 29 · 0.6040 · 17.7 h, Judge result recordedGPT-5.6 Sol · submission 30 · 0.6049 · 18.3 h, Judge result recordedGPT-5.6 Sol · submission 31 · 0.6031 · 18.9 h, Judge result recordedGPT-5.6 Sol · submission 32 · 0.6049 · 19.5 h, Judge result recordedGPT-5.6 Sol · submission 33 · 0.5971 · 20.1 h, Judge result recordedGPT-5.6 Sol · submission 34 · 0.6043 · 20.7 h, Judge result recordedGPT-5.6 Sol · submission 35 · 0.5993 · 21.3 h, Judge result recordedGPT-5.6 Sol · submission 36 · 0.5995 · 21.9 h, Judge result recordedGPT-5.6 Sol · submission 37 · 0.6053 · 22.5 h, Judge result recordedGPT-5.6 Sol · submission 38 · 0.5988 · 23.1 h, Judge result recordedGPT-5.6 Sol · submission 39 · 0.5985 · 23.7 h, Judge result recordedGPT-5.6 Sol · submission 40 · 0.6039 · 24.3 h, Judge result recordedGPT-5.6 Sol · 0.6053Claude Opus 5 · submission 3 · 0.5624024003325582 · 2.176 h, Judge result recorded; below the displayed range, marked at the axis.Claude Opus 5 · submission 4 · 0.31013111490596057 · 2.544 h, Judge result recorded; below the displayed range, marked at the axis.Claude Opus 5 · submission 6 · 0.588914339543402 · 3.782 h, Judge result recorded; below the displayed range, marked at the axis.Claude Opus 5 · submission 17 · 0.571584097878923 · 10.490 h, Judge result recorded; below the displayed range, marked at the axis.Claude Opus 5 · submission 18 · 0.5721456164191883 · 11.101 h, Judge result recorded; below the displayed range, marked at the axis.Claude Opus 5 · submission 21 · 0.5808311487459865 · 12.892 h, Judge result recorded; below the displayed range, marked at the axis.Claude Opus 5 · submission 22 · 0.5616897313679421 · 13.529 h, Judge result recorded; below the displayed range, marked at the axis.Claude Opus 5 · submission 24 · 0.47783659480152735 · 14.775 h, Judge result recorded; below the displayed range, marked at the axis.Claude Opus 5 · submission 25 · 0.2821666630983059 · 15.420 h, Judge result recorded; below the displayed range, marked at the axis.Claude Opus 5 · submission 31 · 0.49152040367063765 · 19.211 h, Judge result recorded; below the displayed range, marked at the axis.Claude Opus 5 · submission 33 · 0.5668735513918998 · 20.447 h, Judge result recorded; below the displayed range, marked at the axis.Claude Opus 5 · submission 36 · 0.5866893299751684 · 22.246 h, Judge result recorded; below the displayed range, marked at the axis.Claude Opus 5 · submission 38 · 0.5769691386941643 · 23.455 h, Judge result recorded; below the displayed range, marked at the axis.GPT-5.6 Sol · submission 3 · 0.5601990639390262 · 2.408 h, Judge result recorded; below the displayed range, marked at the axis.GPT-5.6 Sol · submission 12 · 0.5905210272412961 · 7.866 h, Judge result recorded; below the displayed range, marked at the axis.GPT-5.6 Sol · submission 17 · 0.5735986250297 · 10.799 h, Judge result recorded; below the displayed range, marked at the axis.GPT-5.6 Sol · submission 18 · 0.5533460410634136 · 11.393 h, Judge result recorded; below the displayed range, marked at the axis.GPT-5.6 Sol · submission 19 · 0.3039013154057625 · 11.724 h, Judge result recorded; below the displayed range, marked at the axis.GPT-5.6 Sol · submission 22 · 0.48974277196778726 · 13.490 h, Judge result recorded; below the displayed range, marked at the axis.GPT-5.6 Sol · submission 23 · 0.5574342740981484 · 14.121 h, Judge result recorded; below the displayed range, marked at the axis.GPT-5.6 Sol · submission 24 · 0.5887556372220799 · 14.712 h, Judge result recorded; below the displayed range, marked at the axis.GPT-5.6 Sol · submission 28 · 0.5740402176998849 · 17.108 h, Judge result recorded; below the displayed range, marked at the axis.
submissionrunning bestbaseline · Judge measured▽ below 0.592Time since run start · points mark recorded Judge results

MolmoWeb Interaction Context Allocation

Choose which interaction history, page details, and screenshots a frozen MolmoWeb model sees to improve its next browser-action prediction.

Judged on macro-averaged browser-action prediction quality, scored absolutely.

full task page → task ↗ run log ↗

8 H100 · 24 hPosttrain

Teaser figure from ReasonIR
ReasonIR · source ↗
0.1960.2040.2120.220.228Baseline · 0.21141 reward081624.1Elapsed time (h)Guarded retrieval rewardClaude Opus 5 · submission 2 · 0.21776 · 5.36 h, Judge result recordedClaude Opus 5 · submission 3 · 0.21530 · 12.2 h, Judge result recordedClaude Opus 5 · submission 4 · 0.21161 · 12.8 h, Judge result recordedClaude Opus 5 · submission 5 · 0.21776 · 13.4 h, Judge result recordedClaude Opus 5 · submission 6 · 0.22340 · 17.8 h, Judge result recordedClaude Opus 5 · submission 8 · 0.22340 · 19.1 h, Judge result recordedClaude Opus 5 · submission 9 · 0.20748 · 21.6 h, Judge result recordedClaude Opus 5 · submission 10 · 0.22340 · 22.2 h, Judge result recordedClaude Opus 5 · 0.22340GPT-5.6 Sol · submission 1 · 0.20020 · 3.84 h, Judge result recordedGPT-5.6 Sol · submission 2 · 0.21536 · 6.4 h, Judge result recordedGPT-5.6 Sol · submission 5 · 0.21811 · 14.1 h, Judge result recordedGPT-5.6 Sol · submission 6 · 0.21260 · 16.8 h, Judge result recordedGPT-5.6 Sol · submission 7 · 0.21794 · 19.4 h, Judge result recordedGPT-5.6 Sol · 0.21811Claude Opus 5 · submission 1 · -0.00306 · 2.605 h, Judge result recorded; below the displayed range, marked at the axis.Claude Opus 5 · submission 7 · -0.0049 · 18.421 h, Judge result recorded; below the displayed range, marked at the axis.GPT-5.6 Sol · submission 3 · -0.01126 · 9.015 h, Judge result recorded; below the displayed range, marked at the axis.GPT-5.6 Sol · submission 4 · -0.00368 · 11.532 h, Judge result recorded; below the displayed range, marked at the axis.GPT-5.6 Sol · submission 8 · -0.01372 · 22.018 h, Judge result recorded; below the displayed range, marked at the axis.
submissionrunning bestbaseline · Judge measured▽ below 0.196Time since run start · points mark recorded Judge results

ReasonIR Difficulty Curriculum

Choose how a fixed pool of easy and hard retrieval examples is weighted and ordered to improve ReasonIR-8B, without weakening its general retrieval ability.

Judged on reasoning-retrieval quality, with a penalty if general retrieval regresses, scored absolutely.

full task page → task ↗ run log ↗

Call for Contributors

RSI-Anything: Human–AI collaboration for RSI task contribution.

The agent walks you through every judgment call, flags the pitfalls of task design, and ends with your research as a complete autoresearch environment.

Do you want to challenge frontier agents with your own representative research work?

Do you want to see an agent propose insights you never thought of?

1 hour

All you need is a conversation with our
RSI-Anything Agent.

Any questions? Contact yuetaili@uw.edu, zhuofengli12345@gmail.com, yfeng42@uw.edu

accept pass reject · revise 1 Create a proposal One session with the RSI-Anything Agent about an hour human + agent 2 Submit for automatic review Task Ideas Discussion a review agent returns accept or reject agent 3 Automatic task building the pipeline turns the proposal into a complete autoresearch environment agent 4 Run it and Submit Trajectory private task repository frontier agents take on your research human + agent Contributors receive authorship

Open research community will have the benchmark of our own. Every task here is sourced from academic open-source work, and the credit belongs back with our community.

Call for compute

Compute Partners

Every task runs a real model-development environment on real GPUs. To scale RSI environments for model training, we need more compute.

If you have computation resources to run experiments and want to build exciting, frontier RSI environments together, reach out to us.