go-mlx

History

Snider 19c4823b04 feat(metal): add Llama 3 model support (Llama 3.1 8B validated) Llama shares the Qwen3 loader (same decoder: pre-norm, SwiGLU, GQA). Model type now detected from config.json model_type field instead of weight-only heuristic. Llama 3 chat template and EOS token added. Model tests now clear Metal GPU cache between runs. Llama 3.1 8B Instruct 4-bit: 30 tok/s on M3 Ultra. Co-Authored-By: Virgil <virgil@lethean.io> Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>	2026-02-19 23:06:43 +00:00
..
metal	feat(metal): add Llama 3 model support (Llama 3.1 8B validated)	2026-02-19 23:06:43 +00:00

Snider 19c4823b04 feat(metal): add Llama 3 model support (Llama 3.1 8B validated)

Llama shares the Qwen3 loader (same decoder: pre-norm, SwiGLU, GQA).
Model type now detected from config.json model_type field instead of
weight-only heuristic. Llama 3 chat template and EOS token added.
Model tests now clear Metal GPU cache between runs.

Llama 3.1 8B Instruct 4-bit: 30 tok/s on M3 Ultra.

Co-Authored-By: Virgil <virgil@lethean.io>
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

2026-02-19 23:06:43 +00:00

metal

feat(metal): add Llama 3 model support (Llama 3.1 8B validated)

2026-02-19 23:06:43 +00:00