Open agentic modeling research

Orchard-Agentic

A collection of open-source research projects, shared infrastructure, training recipes, and evaluation resources for building and studying capable agents.

News

  1. [2026-09] AutoEnvScaling — Terminal agents build their own RL environments, a step toward recursive self-improvement: acting as both proposer and solver, Qwen3.6-35B-A3B improves Terminal-Bench 2.1 from 41.1 to 53.3. Project page ↗

  2. [2026-09] SecureVibe — A unified SFT, RL, and on-policy distillation stack for training security-aware coding agents. Project page ↗ · Code ↗

  3. [2026-09] SecureVibeEval — A shared evaluation suite for endpoint models and coding CLIs across SecureGen, AutoBax, BaxBench, and SusVibes. Code ↗

  4. [2026-09] ExecCritic — Learned tests and feedback-guided repair reach 72.6% on SWE-bench Verified, an 11.4-point gain over the no-test baseline. Paper ↗ Code ↗

  5. [2026-08] TRACE — Turn-level credit assignment improves BrowseComp-Plus from 7.2 to 35.6 at 4B and 8.4 to 42.6 at 30B. Paper ↗ Code ↗

  6. [2026-07] OpenForge RL — Harness-native training reaches 37.7 on OSWorld-Verified and 72.3 on WebVoyager while reducing the train–deploy mismatch. Paper ↗

  7. [2026-06] OpenWebRL — A 4B visual agent trained on live websites reaches 67.0% on Online-Mind2Web and 64.0% on DeepShop. Paper ↗

  8. [2026-05] Orchard — A reusable environment layer supports open agent training across coding, GUI, and assistant workflows. Paper ↗ Project page ↗

Projects

2026-09 · Environment generation

AutoEnvScaling

Automating the data flywheel with terminal agents: a proposer builds RL environments from the solver's rollouts, and one model co-evolves in both roles toward recursive self-improvement.

Topics: Environment generation · Recursive self-improvement · Terminal agents

2026-09 · Security-aware training

SecureVibe

Training workflows for security-aware coding agents using supervised fine-tuning, reinforcement learning, and on-policy distillation.

Topics: Secure coding · SFT · RL · On-policy distillation

2026-09 · Security evaluation

SecureVibeEval

Evaluation tooling for security-aware coding agents on SecureGen, AutoBax, BaxBench, and SusVibes, including endpoint runners, graders, and a multi-CLI harness.

Topics: Agent evaluation · Security benchmarks · CLI harnesses

2026-09 · SWE training

execcritic

A behavior-contract generated-test, self-repair, and SWE-RL training kit with isolated generated-test gates and official SWE-bench verification.

Topics: Behavior contracts · Self-repair · SWE-RL

2026-08 · Credit assignment

TRACE

Turn-level reward assignment via credit estimation for long-horizon agents, assigning credit at tool-call boundaries.

Topics: Turn-level rewards · Credit estimation · Long-horizon agents

2026-07 · Harness-native training

OpenForge RL

Training agents inside their real deployment harnesses to reduce train-deploy mismatch across coding, GUI, and assistant workflows.

Topics: Agent harnesses · Reinforcement learning · Deployment

2026-06 · Live web agents

OpenWebRL

Online, multi-turn reinforcement learning for agents operating on live websites with fault-tolerant browser interaction.

Topics: Online RL · Live web · Computer use

2026-05 · Foundational project

Orchard

An open-source framework for scalable agentic modeling across software engineering, computer use, and personal-assistant tasks.

Artifacts: Orchard Env · Orchard-SWE · Orchard-GUI · Orchard-Claw

Shared infrastructure

Orchard Env is a thin, Kubernetes-native environment service for isolated, multi-turn agent interaction. It remains independent from any one trainer, harness, or task domain.

View Orchard Env on GitHub ↗
  • Sandbox lifecycle Isolated environments with resource limits, timeouts, and network policy.
  • Multi-turn execution Commands, files, patches, and persistent state throughout an agent trajectory.
  • Harness agnostic The same primitives across different trainers, harnesses, and inference backends.
  • Built for scale Concurrent Kubernetes sandboxes with a low-latency in-pod execution path.