Text Generation
Transformers
Safetensors
qwen3_5_moe_text
qwen3.5
qwen
terminal-agent
terminus-2
sft
conversational

Qwen3.5-35B-A3B terminal agent — step 5,271

This is the step-5,271 intermediate checkpoint of the terminal agent run. It is a supervised fine-tune of Qwen/Qwen3.5-35B-A3B-Base for agentic text generation. Terminal-agent SFT over Terminus-2-focused trajectories from two modern teacher lineages.

This repository contains BF16 Hugging Face safetensors exported from the finalized Megatron distributed checkpoint. It is not quantized.

Checkpoint identity

Field Value
Hugging Face repository eewer/qwen3.5-35b-a3b-base-terminal-agent-step5271
Run name terminal
Training iteration 5,271
Planned iterations / epoch 5,271
Position in planned epoch 100.00%
Base model Qwen/Qwen3.5-35B-A3B-Base
Source checkpoint /KRAFTON/WORKSPACE/wbl-workspace/posttraining-2606/areal_runs/qwen35_b300/sft/combined-8gpu-tp1-ep8-cp1-mbs1-full-router1-shared0-gbs32-lr2e5/checkpoints/iter_0005271
Export format BF16 Hugging Face safetensors
Context length used for SFT 32,768 tokens

Training configuration

  • Hardware: one node with 8 NVIDIA B300 GPUs.
  • Framework: Megatron-Bridge / Megatron-Core in the NVIDIA NeMo 26.06 environment.
  • Parallelism: TP=1, PP=1, CP=1, EP=8, DP=8.
  • Micro batch size: 1 per data-parallel rank.
  • Global batch size: 32.
  • Activation recomputation: full, uniform recomputation.
  • MoE router fusion: enabled; shared-expert overlap, gradient-reduce overlap, and parameter-gather overlap disabled.
  • Precision: BF16 training/model tensors with the recipe's precision-aware optimizer state; exported weights are BF16.
  • Learning-rate schedule: cosine decay from a peak of 2e-5 to 2e-6.
  • Warmup: 158 optimizer steps for every run.
  • Checkpoint interval: 500 optimizer steps.
  • Sequence packing: No; each trajectory remains an independent 32K sample.
  • Objective: assistant-only causal language modeling. System, user, and tool messages are context rather than prediction targets.

Training data

The materialized source mixture has 168,646 rows; 168,646 rows enter training after the 32K boundary check. The complete training representation contains 2,912,501,244 effective/rendered tokens.

Source dataset Selected rows
nvidia/Nemotron-Terminal-Corpus 79,201
open-thoughts/OpenThoughts-Agent-SFT-100K 89,445

The dataset is a deterministic shuffled union. Trajectory boundaries are preserved (no sequence packing). System and user tokens are context-only. Valid assistant turns are supervised; known malformed Nemotron recovery turns carry a per-turn zero loss mask while the surrounding valid assistant turns remain trainable.

The tokenizer is from Qwen3.5, with a Qwen3.6 thinking-preserving chat template. The template keeps reasoning content in assistant messages and preserves tool declarations when the source supplies a validated schema.

Intended use

This checkpoint is intended for research on terminal, software-engineering, tool-use, and interactive agents. Use the included chat template and supply tool schemas expected by the target harness. It can be served with Transformers-compatible runtimes that support the Qwen3.5 MoE architecture.

This is an intermediate training checkpoint, not a polished instruction model. Compare checkpoints using held-out loss and task-level agent evaluations before selecting one for deployment.

Limitations and safety

  • No complete benchmark suite is claimed in this model card.
  • Agent trajectories can produce destructive shell commands, modify files, call tools, or expose secrets. Run the model in an isolated environment with scoped credentials.
  • Tool names and schemas vary across source harnesses. A caller must provide the schema appropriate for its own environment; do not assume a generated call is valid or safe.
  • Dataset filtering and deduplication reduce, but cannot guarantee removal of, benchmark contamination, incorrect reasoning, insecure code, or teacher-model artifacts.
  • The model can hallucinate successful tool execution and should not be trusted without checking actual environment observations.
  • This checkpoint inherits the base model's limitations and license.

Reproducibility notes

All three runs use the same B300 execution topology, optimizer schedule, 158-step warmup, and checkpoint cadence. The terminal_swe and terminal_swe_general runs differ from the terminal run primarily in their source mixture and use of sequence packing. The mixed datasets were validated for unique conversation hashes, deterministic shuffle order, non-empty reasoning in every assistant turn, maximum rendered length, and tool-call/schema consistency.

Downloads last month
20
Safetensors
Model size
35B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for eewer/qwen3.5-35b-a3b-base-terminal-agent-step5271

Finetuned
(70)
this model

Datasets used to train eewer/qwen3.5-35b-a3b-base-terminal-agent-step5271