core/go-mlx: Native Apple Metal GPU inference via mlx-c bindings

Native Apple Metal GPU inference via mlx-c bindings

Find a file

Snider 5644857034 feat(metal): implement batch inference (Classify, BatchGenerate) - Add ForwardMasked to InternalModel, Gemma3 and Qwen3 architectures - Thread attention mask through decoder layers and SDPA calls - Use ScaledDotProductAttentionWithMask when explicit mask provided - Create batch.go with padded batching, mask construction, Classify (prefill-only) and BatchGenerate (autoregressive) implementations - Wire Classify/BatchGenerate through metalAdapter to go-inference - Tests: mask unit tests (shape, values, multi-batch), Classify with 4 prompts (152 prompts/s), WithLogits, BatchGenerate with 2 prompts Co-Authored-By: Virgil <virgil@lethean.io> Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>		2026-02-19 23:28:15 +00:00
cpp	fix(metal): address 4 minor code review items	2026-02-19 21:36:40 +00:00
docs/plans	docs: batch inference API design (Phase 5)	2026-02-19 23:18:38 +00:00
internal/metal	feat(metal): implement batch inference (Classify, BatchGenerate)	2026-02-19 23:28:15 +00:00
.gitignore	chore: gitignore dist/ (CMake install output)	2026-02-19 19:30:23 +00:00
CLAUDE.md	feat(api): migrate to go-inference shared interfaces	2026-02-19 20:15:42 +00:00
CMakeLists.txt	feat: extract go-mlx from go-ai as standalone Metal inference package	2026-02-19 17:57:37 +00:00
FINDINGS.md	fix(metal): address 4 minor code review items	2026-02-19 21:36:40 +00:00
go.mod	feat(api): migrate to go-inference shared interfaces	2026-02-19 20:15:42 +00:00
mlx.go	feat(api): migrate to go-inference shared interfaces	2026-02-19 20:15:42 +00:00
mlx_stub.go	feat: extract go-mlx from go-ai as standalone Metal inference package	2026-02-19 17:57:37 +00:00
mlx_test.go	feat(metal): implement batch inference (Classify, BatchGenerate)	2026-02-19 23:28:15 +00:00
register_metal.go	feat(metal): implement batch inference (Classify, BatchGenerate)	2026-02-19 23:28:15 +00:00
TODO.md	feat(metal): add mixed precision training via LoRAConfig.DType (Phase 3)	2026-02-19 23:13:49 +00:00