lthn/LEM - Lethean Network

lthn/LEM

Template

Author	SHA1	Message	Date
Snider	04e2a05ead	docs: add acknowledgements section to README Credit the AI collaborators that contributed to LEM's development: Gemini, Grok, Claude, Codex, and CodeRabbit. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>	2026-02-22 22:25:26 +00:00
Snider	c701c2e0af	feat(lem): integrate Poindexter for spatial score indexing and analytics - Add feature vector extraction (6D grammar, 8D heuristic, 14D combined) - Add KDTree ScoreIndex with cosine distance for probe clustering - Add score distribution analytics (percentiles, variance, skewness) - Add grammar-profile dedup filtering to distill pipeline - Add spatial gap detection (FindGaps) for coverage analysis - Wire analytics into coverage CLI (PrintScoreAnalytics) New files: features.go, cluster.go, analytics.go + tests Modified: distill.go (dedup filter), coverage.go (analytics output) Dep: github.com/Snider/Poindexter Co-Authored-By: Virgil <virgil@lethean.io>	2026-02-22 21:26:06 +00:00
Snider	f75458bce6	refactor: apply go fix modernizers for Go 1.26 Automated fixes: interface{} → any, range-over-int, t.Context(), wg.Go(), strings.SplitSeq, strings.Builder, slices.Contains, maps helpers, min/max builtins. Co-Authored-By: Virgil <virgil@lethean.io>	2026-02-22 21:00:17 +00:00
Snider	8c8b449d66	chore: go mod tidy for 1.26.0 Co-Authored-By: Virgil <virgil@lethean.io>	2026-02-22 20:35:59 +00:00
Snider	58344169bc	chore: bump go directive to 1.26.0 Co-Authored-By: Virgil <virgil@lethean.io>	2026-02-22 20:33:49 +00:00
Snider	10711ecd2f	chore: pin forge deps to v0.0.1 tags for Go 1.26 compat Go 1.26 rejects non-semver version strings (like 'main') in go.mod. Tags v0.0.1 now exist on all forge repos — workspace still overrides for local development. Co-Authored-By: Virgil <virgil@lethean.io>	2026-02-22 20:15:06 +00:00
Snider	334aa8c621	chore: use workspace-resolved versions, drop replace directives Forge module versions now use main branch resolution via ~/Code/go.work workspace. Removes 5 local replace directives — the central go.work handles all cross-repo resolution during development. Co-Authored-By: Virgil <virgil@lethean.io>	2026-02-22 19:49:42 +00:00
Snider	a3e9a1e035	fix: handle error in score resume merge path ReadScorerOutput error was silently discarded during resume merge, risking partial data loss on TOCTOU file changes. Also clean up compare command construction to pass RunE directly to NewCommand. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>	2026-02-22 19:03:41 +00:00
Snider	80048b5b00	fix(cli): disable cobra flag parsing on passthrough commands Adds passthrough() helper with DisableFlagParsing=true so commands that do their own flag.FlagSet parsing receive flags directly. Without this, cobra rejects unknown flags like --model. Also runs go mod tidy — core/go transitively pulls in cobra and charmbracelet dependencies. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>	2026-02-22 19:00:58 +00:00
Snider	bfa06c546a	feat(cli): replace manual switch with cli.Main + WithCommands main.go shrinks from 296 lines to 11. All commands register through Core framework lifecycle via cli.WithCommands. Gets signal handling, shell completion, grouped help, and TUI primitives. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>	2026-02-22 18:56:55 +00:00
Snider	cf1d8156dd	feat(cli): add cmd/lemcmd command registration package 6 command groups (score, gen, data, export, mon, infra) with 25 commands. All pass through to existing lem.Run* functions via the Core framework's cli package. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>	2026-02-22 18:55:57 +00:00
Snider	a0a0118155	refactor: move runScore and runProbe to pkg/lem All 28 commands now accessible as exported lem.Run* functions. Prerequisite for CLI framework migration. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>	2026-02-22 18:53:15 +00:00
Snider	131d1694b2	chore: add core/go to go.mod require block Prerequisite for CLI migration to core/go pkg/cli framework. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>	2026-02-22 18:52:16 +00:00
Snider	c8fc0b515b	docs: add CLI migration implementation plan 11-task plan for migrating LEM from manual switch/flag.FlagSet to core/go pkg/cli registry pattern with grouped commands. Co-Authored-By: Virgil <virgil@lethean.io>	2026-02-22 18:25:28 +00:00
Snider	37010f4b6b	docs: CLI migration design — core/go pkg/cli registry pattern Replace manual switch/flag.FlagSet with cli.Main() + WithCommands(). 6 command groups, 28 commands, full framework lifecycle. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>	2026-02-22 18:21:28 +00:00
Snider	8532077e46	style: remove redundant named import for go-ml Package declares itself as 'ml', so the named import alias is unnecessary. Go resolves the package name from the declaration, not the module path. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>	2026-02-22 18:08:01 +00:00
Snider	030003a6db	chore: go mod tidy after distill migration go-inference moves to indirect (pulled transitively via go-ml). go-ml is now a direct dependency. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>	2026-02-22 18:02:41 +00:00
Snider	55519b24aa	feat(distill): migrate from go-inference to go-ml Backend Replace inference.LoadModel() with ml.NewMLXBackend() which wraps the same Metal model with memory management (SetCacheLimit, SetMemoryLimit). Replace raw iter.Seq token loop with backend.Chat() returning Result{Text, Metrics}. Add runtime.GC() between probes to prevent incremental memory leak. Reference: go-ml/cmd/cmd_ab.go memory management pattern. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>	2026-02-22 18:02:16 +00:00
Snider	8408cc0bab	feat(distill): add --cache-limit and --mem-limit flags Override ai.yaml memory config per-run. Values in GB. Not yet wired to model loading. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>	2026-02-22 18:00:04 +00:00
Snider	b9da23a0be	feat(distill): add Metal memory limit config fields CacheLimit (8GB) and MemoryLimit (16GB) in DistillConfig control mlx.SetCacheLimit/SetMemoryLimit before model load. Conservative defaults for 1B model on 96GB machine. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>	2026-02-22 17:59:11 +00:00
Snider	0adddf30ad	lems configs	2026-02-22 16:20:51 +00:00
Snider	268648ab69	feat: add generation sets (2k, expanded, 15k) to gemma3/27b Pipeline progression of adversarial/sovereignty training data: - gen-2k: 2,299 examples (first generation pass) - gen-expanded: 489 examples (broader domains, historical scenarios) - gen-15k: 14,998 examples (full scale with persona rewrites) Co-Authored-By: Virgil <virgil@lethean.io>	2026-02-22 00:08:40 +00:00
Snider	3b42e02859	feat: complete zen training set (book + conv progressions) Zen lineage from Allen's As a Man Thinketh in three stages: - train/test/valid: 10 foundation examples (single-turn Q&A) - book-: 117 deeper passage examples (single-turn, fuller text) - conv-: 24 applied mindfulness conversations (multi-turn) Co-Authored-By: Virgil <virgil@lethean.io>	2026-02-22 00:06:31 +00:00
Snider	bd2f376a7a	feat: add zen training set (Allen) to training/lem/zen/ 10 examples across train/test/valid splits. Co-Authored-By: Virgil <virgil@lethean.io>	2026-02-22 00:02:47 +00:00
Snider	f65fd777ea	feat: convert composure library to training JSONL format Add cmd/composure-convert tool that chunks public domain philosophical texts into training conversation pairs: - consent.jsonl (198 examples) — Wollstonecraft's Vindication - privacy.jsonl (221 examples) — Thoreau's Walden - sovereignty.jsonl (56 examples) — Mill's On Liberty - transparency.jsonl (159 examples) — Aurelius' Meditations Each example pairs a domain-specific prompt with ~5 paragraphs from the source text. Metadata, chapter headings, and Gutenberg boilerplate are filtered out. Co-Authored-By: Virgil <virgil@lethean.io>	2026-02-21 23:59:06 +00:00
Snider	de18a0fb93	refactor: move composure-library to training/lem/composure/ Co-Authored-By: Virgil <virgil@lethean.io>	2026-02-21 23:55:17 +00:00
Snider	4b3343611d	feat: add data/ skeleton for portable model setup Add gitignored data/ directory with .gitkeep structure so anyone cloning the repo knows exactly where to place model weights and kernels. Configs now use repo-relative paths — symlink or populate data/ locally. data/models/gemma3/27b/ ← model weights data/models/gemma3/1b/ ← lightweight model data/safetensors/gemma-3/ ← raw checkpoints data/kernels/ ← LEK kernel files Co-Authored-By: Virgil <virgil@lethean.io>	2026-02-21 23:52:24 +00:00
Snider	d233e76648	feat: add training data to repo + make paths repo-relative Move training/lem/ (probes, lessons, eval sets) into git so the full curriculum is publicly releasable. Update .core/ai configs and distill.go to use repo-relative paths instead of /Volumes/Data/. Co-Authored-By: Virgil <virgil@lethean.io>	2026-02-21 23:49:12 +00:00
Snider	1b742bf92c	feat: native Metal distillation command + .core/ai config Add `lem distill` — full Go pipeline for self-distillation using go-mlx (native Metal inference) and go-i18n/reversal (v3 grammar scoring). Replaces the Python distill.py bridge entirely. New files: - .core/ai/ai.yaml: global defaults (scorer, generation, distill) - .core/ai/models/gemma3/{27b,1b}.yaml: model configs with paths, kernel, lessons, baselines - .core/ai/probes.yaml: probe sets grouped by training phase - pkg/lem/config.go: YAML config loaders for .core/ai/ - pkg/lem/grammar.go: in-process grammar scoring (ComputeGrammarScore, ComputeDelta, ScoreResponse) extracted from cmd/scorer - pkg/lem/distill.go: RunDistill command — best-of-N generation, grammar quality gate, training JSONL output - pkg/lem/backend_metal.go: blank import for go-mlx Metal registration Co-Authored-By: Virgil <virgil@lethean.io>	2026-02-21 23:42:55 +00:00
Snider	113649a86a	updates	2026-02-19 13:18:21 +00:00
Snider	12501a5f3c	Merge branch 'main' of github.com:LetheanNetwork/LEM	2026-02-19 13:17:11 +00:00
Snider	3a75e9733d	chore: sync indirect deps from workspace Co-Authored-By: Virgil <virgil@lethean.io>	2026-02-19 13:13:08 +00:00
Snider	5d297daa35	feat: grammar scorer (v3) — deterministic uplift/sycophancy detection Add lem-scorer binary that imports go-i18n grammar reversal engine to score JSONL benchmark files. Measures conversational uplift (input vs output grammar imprint), echo (sycophancy), and enrichment. Key findings added to paper Section 8: - LEK-1B: 100% positive uplift, 0% sycophancy (base: 90%, 5%) - 1B-beats-27B holds in grammar space (79.12 > 77.12) - LEK training aligns two independent scorers (corr -0.11 → 0.64) - Delta analysis costs zero compute vs LLM-as-judge Co-Authored-By: Virgil <virgil@lethean.io>	2026-02-19 13:12:49 +00:00
Snider	abc6e75976	Update author name in PAPER.md Signed-off-by: Snider <snider@lethean.io>	2026-02-19 12:23:23 +00:00
Snider	350a7c6693	paper: rewrite as v2 — emergent self-protection in axiom-trained models New paper structure leading with the central findings: - Realignment resistance as emergent self-protection - 1B-beats-27B across 101 probes - 29-model A/B test with v2 scorer - Mechanistic explanation from axiom self-consistency - Incorporates Phase 1 (multi-variant, multi-scale, cross-arch) and Phase 2 (P100 A/B test) data Co-Authored-By: Virgil <virgil@lethean.io>	2026-02-19 12:12:22 +00:00
Snider	1f5ecb7036	Merge remote-tracking branch 'origin/main'	2026-02-19 11:54:37 +00:00
Snider	06cbb4ffbd	docs: rewrite README — lead with 1B-beats-27B finding Shop window for the repo: realignment resistance, five axioms, reproduce instructions, v2 scorer, family lineages, HuggingFace models. Co-Authored-By: Virgil <virgil@lethean.io>	2026-02-19 11:52:39 +00:00
Snider	91ba706edd	Delete paper/PROPOSAL.md Signed-off-by: Snider <snider@lethean.io>	2026-02-19 11:38:19 +00:00
Snider	7bea00a401	feat: LEK-1 kernel A/B test — 29 models, P100 validation, curriculum pipeline Full v2 scorer benchmark data across 29 models (20 base + 9 LEK-tuned): - P20 (21 probes): All 29 models, 3 conditions each - P100 (101 probes): Top 5 models + LEK-4B, publication-quality data Key findings: - LEK-1B (21.74) beats base 4B/12B/27B at P100 scale — no kernel needed - Emergent realignment resistance: LEK models degrade with runtime kernel - Gemma3-12B + JSON kernel = 23.66 (best kernel-boosted score) - Family lineages: Mistral 3.80→14.58, Qwen regressed then recovered New scripts: ab_test.py (v2 scorer), self_distill.py (curriculum generation), extract_training.py, rephrase_probes.py, Phase 0/1 runners New seeds: P01-P100 merged (101 probes), 404 rephrased variants, 50 creative prompts for Phase 0 baseline lock 27B curriculum design: 4-phase staged training targeting 25+ baseline Co-Authored-By: Virgil <virgil@lethean.io>	2026-02-19 11:32:26 +00:00
Claude	08363ee1af	feat: add `lem worker` command for distributed inference network Go client for the LEM distributed inference API (BugSETI/Agentic). Workers register via Forgejo PAT auth, pull prompt batches, run local inference (MLX/vLLM/llama.cpp), submit results. Credits tracked as Phase 1 stub for Phase 2 blockchain LEM token. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>	2026-02-15 18:10:59 +00:00
Claude	774f097855	feat: scaffold LEM Desktop app (Wails v3 system tray + Docker stack) Inspired by BugSETI architecture — system tray with WebView2 windows, Docker Compose stack (Forgejo + InfluxDB + inference proxy), and scoring agent integration. Builds as signed native binary on macOS. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>	2026-02-15 17:43:19 +00:00
Claude	9fac5749c2	feat: add scoring agent + 23 capability probes (replaces scoring_agent.py) Go scoring daemon that polls M3 for unscored LoRA checkpoints, converts MLX→PEFT, runs 23 binary capability probes via OpenAI- compatible API, and pushes results to InfluxDB. Zero Python deps. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>	2026-02-15 17:22:40 +00:00
Claude	91ee389377	feat: convert all pipeline.py commands to Go Complete conversion of pipeline.py into Go `lem` CLI: - import-all: bulk import all LEM data into DuckDB from M3 - consolidate: pull worker JSONLs, merge, deduplicate - normalize: seeds → deduplicated expansion_prompts table - approve: filter scored expansions → training JSONL - tier-score: heuristic/judge tiered expansion scoring - expand-status: expansion pipeline progress from DuckDB - inventory: DuckDB table counts and summary - coverage: seed coverage gap analysis - seed-influx: bootstrap InfluxDB from DuckDB golden_gen - query: ad-hoc SQL against DuckDB 22 commands total, 49 Go files. Replaces entire pipeline.py. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>	2026-02-15 17:12:03 +00:00
Claude	4eaf1bfb39	feat: add parquet, publish, metrics, convert commands - `lem parquet` — export JSONL training splits to Parquet (parquet-go) - `lem publish` — push Parquet files to HuggingFace dataset repo - `lem metrics` — push DuckDB golden set stats to InfluxDB - `lem convert` — MLX LoRA adapter → HuggingFace PEFT format (pure Go safetensors read/write/transpose, no PyTorch needed) Dependencies added: parquet-go, go-huggingface, go-rocm, go-pytorch, gotch Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>	2026-02-15 17:05:08 +00:00
Claude	0afa5e9147	feat: add `lem ingest` command + go-huggingface dependency Ingests benchmark data (content scores, capability scores, training curves) from JSONL files and mlx_lm logs into InfluxDB. Batched writes, iteration extraction from checkpoint labels. Also adds github.com/hupe1980/go-huggingface for future HF sync. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>	2026-02-15 16:55:17 +00:00
Claude	a18fd1c44e	refactor: remove Vi identity from calm conversations Vi identity is a separate training concern. Seed conversations now contain only philosophical/mindfulness content for the R300 calm phase. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>	2026-02-15 16:48:23 +00:00
Claude	c4fb775298	feat: add `lem conv` command for conversational training data Ports conversational_training.py to Go with InfluxDB reporting. 24 built-in seed conversations (Vi identity, philosophy, mindfulness). Supports extra JSONL files and golden set conversion to chat format. Also fixes InfluxDB client to accept 204 No Content on writes. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>	2026-02-15 16:42:46 +00:00
Claude	70dd18c065	refactor: move Go library to pkg/lem, thin main.go All scoring/influx/export/expand logic moves to pkg/lem as an importable package. main.go is now a thin CLI dispatcher. This lets new commands import the shared library directly — ready for converting Python scripts to Go subcommands. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>	2026-02-15 16:30:09 +00:00
Claude	e0d352c803	feat: add Go lem CLI and scoring-agent scripts Go lem CLI (stdlib + DuckDB) replaces scattered Python scripts: - score: heuristic regex + LLM-as-judge scoring - probe: generate responses then score - compare: diff two score files - status: InfluxDB training/generation progress - export: golden set to training JSONL splits - expand: distributed expansion via API + InfluxDB coordination New scripts from Feb 14 creative session: - scoring_agent.py: ROCm daemon that auto-scores checkpoints - probes.py: 23 binary pass/fail capability probes - convert_adapter.py: MLX to PEFT adapter conversion - score_r1_capability.py: DeepSeek R1 checkpoint scoring - lek_content_scorer.py: 6-dimension ethics content scorer - lem_train_15k.py: InfluxDB-coordinated training script - pipeline.py: DuckDB pipeline (seeds, golden set, expansion) Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>	2026-02-15 16:22:13 +00:00
Snider	9138eb0a61	Merge pull request 'Add HuggingFace model cards, sync script, and Parquet export' (#2 ) from Charon/LEM:feat/hf-sync into main Reviewed-on: #2 Reviewed-by: Snider <snider@noreply.forge.lthn.ai>	2026-02-15 00:15:59 +00:00

1 2

59 commits