Coding Agent: MoE¶
| GPU | Model | Actor | Rollout | Router replay |
|---|---|---|---|---|
| 4× B200 | Qwen/Qwen3.5-35B-A3B |
Megatron | vLLM | R3 |
This is the MoE variant of the Coding Agent example. It reuses the same
SWE-smith data, Kubernetes controller, repository images, agent, and reward. The existing
Qwen/Qwen3.5-9B FSDP path remains available through examples/swe_smith/run.sh.
Results¶
With pure RL, Qwen3.5-35B-A3B improves from 47.8% to 61.6% on SWE-bench Verified after 1,792 training examples (about 1.8K), a gain of 13.8 percentage points. Results use the official SWE-bench Verified harness on all 500 instances; trained example counts assume 16 examples per step.
| Step | Trained examples | Result on SWE-bench Verified |
|---|---|---|
| 0 (base) | 0 | 47.8% (239/500) |
| 64 | 1,024 | 58.8% (294/500) |
| 112 | 1,792 | 61.6% (308/500) |
| 176 | 2,816 | 60.6% (303/500) |
Why R3¶
An MoE token can select different experts during rollout and training. R3 records vLLM's rollout
routing and passes it through the Agent Lightning event, triplet, and DataProto pipeline so the
Megatron actor update replays the same experts.
The MoE launcher enables both sides:
actor_rollout_ref.rollout.enable_rollout_routing_replay=true
actor_rollout_ref.actor.megatron.router_replay.mode=R3
Run¶
Use the MoE wrapper for all three roles:
# Machine B
export AGL_SERVER_PUBLIC_HOST=<address-reachable-from-controller-and-pods>
export AGL_KEY=<shared-secret>
examples/swe_smith/run_moe.sh server
# Machine A
export AGL_SERVER_PUBLIC_HOST=<gateway-address>
export AGL_KEY=<same-shared-secret>
export AGL_NAMESPACE=agents
examples/swe_smith/run_moe.sh controller
# Machine B
export AGL_KEY=<same-shared-secret>
examples/swe_smith/run_moe.sh trainer
run_moe.sh selects train_smith_agent_moe.py, requests routed experts from the Gateway, and keeps
the standard 9B launcher unchanged. R3 route compaction requires the default trajectory aggregator;
do not override it to transition. Other additional arguments are VERL dotlist overrides:
examples/swe_smith/run_moe.sh trainer \
trainer.total_training_steps=100 \
actor_rollout_ref.rollout.n=4
The default topology is PP=1, TP=1, EP=4, and ETP=1 with parameter, optimizer, and gradient offload. Treat it as a four-B200 starting point and tune batch sizes for the available memory.