Skip to content

Latest commit

 

History

History
99 lines (72 loc) · 3.35 KB

File metadata and controls

99 lines (72 loc) · 3.35 KB

Model Selection

TraceOtter should optimize for trainability, licensing, long-horizon coding behavior, and future RL/evaluator compatibility. Raw leaderboard rank is not enough.

Default

Qwen/Qwen3-4B-Instruct-2507

Why this is the default:

  • It is small enough to fine-tune first.
  • LLaMA-Factory documents Qwen3 SFT support with template: qwen3_nothink.
  • It keeps the project on the newer Qwen3 tokenizer/template path instead of the older Qwen2.5-Coder path.
  • A 4B dense model is easier to debug, evaluate, quantize, and serve than a larger MoE model.

Generated LLaMA-Factory settings:

model_name_or_path: Qwen/Qwen3-4B-Instruct-2507
template: qwen3_nothink
trust_remote_code: true
cutoff_len: 8192
finetuning_type: lora

Upgrade Target

Qwen/Qwen3-Coder-30B-A3B-Instruct

Use this when GPU budget allows and the evaluator is stable. It is a stronger agentic-coding target, but its total parameter footprint makes iteration slower. For TraceOtter, it should be the second-stage model, not the first bootstrap model.

Teacher Models (for distillation)

On-policy distillation (roadmap M6) and richer teacher-graded targets (M5) need a teacher — a stronger model that grades the student's tokens. The teacher only needs to expose token log-probabilities (compute_logprobs); it is never shipped. Reasonable open teachers as of mid-2026:

Qwen/Qwen3-Coder-480B-A35B-Instruct   # strong open coding teacher
GLM-5.2 / DeepSeek V4 / Kimi K2.6      # alternative strong open teachers

The frontier model that produced the original Codex/Claude trajectories is also a valid teacher — those logs are already teacher rollouts.

Trainer Per Stage

Match the tool to the training stage. Do not assume one trainer covers all of them (see docs/roadmap/RESEARCH-2026.md):

Stage Tool Notes
SFT / DPO / KTO / reward modeling LLaMA-Factory TraceOtter's first target; Qwen3 templates supported.
On-policy distillation (M6) TRL distillation trainer / Tinker cookbook Reverse-KL per-token; not a first-class LLaMA-Factory feature.
Agentic RL — GRPO/DAPO (M7) verl Production RL; replaces the older Agent-R1 framing.

Watchlist

Qwen/Qwen3-Coder-Next   # Apache-2.0, built for coding agents + local dev

Qwen3-Coder-Next (released early 2026) is the leading candidate to become the next default student: permissive license, coding-agent focus, and small enough for local iteration. Promote it only after the selection rule below passes.

Newer is not automatically better for this project; unsupported templates and inference backends can waste days. Validate template support and a smoke fine-tune before changing the default.

Why Not Qwen2.5-Coder Anymore

Qwen2.5-Coder was a reasonable initial default because it was small and easy to train. It is now the wrong long-term anchor because TraceOtter needs newer Qwen3-compatible templates, longer trajectory windows, and a path toward agentic-code models.

Selection Rule

Promote a new default only when all of these are true:

  • LLaMA-Factory has a documented or locally verified template.
  • The model has a permissive enough license for OSS and private commercial use.
  • A 10-file local smoke pipeline exports non-empty SFT data.
  • A one-epoch LoRA smoke run completes.
  • Held-out route/verification evaluation does not regress.