AutoEnvScaling
Automating the data flywheel with terminal agents: a proposer builds RL environments from the solver's rollouts, and one model co-evolves in both roles toward recursive self-improvement.
Open agentic modeling research
A collection of open-source research projects, shared infrastructure, training recipes, and evaluation resources for building and studying capable agents.
[2026-09] AutoEnvScaling — Terminal agents build their own RL environments, a step toward recursive self-improvement: acting as both proposer and solver, Qwen3.6-35B-A3B improves Terminal-Bench 2.1 from 41.1 to 53.3. Project page ↗
[2026-09] SecureVibe — A unified SFT, RL, and on-policy distillation stack for training security-aware coding agents. Project page ↗ · Code ↗
[2026-09] SecureVibeEval — A shared evaluation suite for endpoint models and coding CLIs across SecureGen, AutoBax, BaxBench, and SusVibes. Code ↗
[2026-09] ExecCritic — Learned tests and feedback-guided repair reach 72.6% on SWE-bench Verified, an 11.4-point gain over the no-test baseline. Paper ↗ Code ↗
[2026-08] TRACE — Turn-level credit assignment improves BrowseComp-Plus from 7.2 to 35.6 at 4B and 8.4 to 42.6 at 30B. Paper ↗ Code ↗
[2026-07] OpenForge RL — Harness-native training reaches 37.7 on OSWorld-Verified and 72.3 on WebVoyager while reducing the train–deploy mismatch. Paper ↗
[2026-06] OpenWebRL — A 4B visual agent trained on live websites reaches 67.0% on Online-Mind2Web and 64.0% on DeepShop. Paper ↗
[2026-05] Orchard — A reusable environment layer supports open agent training across coding, GUI, and assistant workflows. Paper ↗ Project page ↗
Automating the data flywheel with terminal agents: a proposer builds RL environments from the solver's rollouts, and one model co-evolves in both roles toward recursive self-improvement.
Training workflows for security-aware coding agents using supervised fine-tuning, reinforcement learning, and on-policy distillation.
Evaluation tooling for security-aware coding agents on SecureGen, AutoBax, BaxBench, and SusVibes, including endpoint runners, graders, and a multi-CLI harness.
A behavior-contract generated-test, self-repair, and SWE-RL training kit with isolated generated-test gates and official SWE-bench verification.
Turn-level reward assignment via credit estimation for long-horizon agents, assigning credit at tool-call boundaries.
Training agents inside their real deployment harnesses to reduce train-deploy mismatch across coding, GUI, and assistant workflows.
Online, multi-turn reinforcement learning for agents operating on live websites with fault-tolerant browser interaction.
An open-source framework for scalable agentic modeling across software engineering, computer use, and personal-assistant tasks.
Orchard Env is a thin, Kubernetes-native environment service for isolated, multi-turn agent interaction. It remains independent from any one trainer, harness, or task domain.
View Orchard Env on GitHub ↗