train-kit / README.md
  1
  2
  3
  4
  5
  6
  7
  8
  9
 10
 11
 12
 13
 14
 15
 16
 17
 18
 19
 20
 21
 22
 23
 24
 25
 26
 27
 28
 29
 30
 31
 32
 33
 34
 35
 36
 37
 38
 39
 40
 41
 42
 43
 44
 45
 46
 47
 48
 49
 50
 51
 52
 53
 54
 55
 56
 57
 58
 59
 60
 61
 62
 63
 64
 65
 66
 67
 68
 69
 70
 71
 72
 73
 74
 75
 76
 77
 78
 79
 80
 81
 82
 83
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
# ardegazu-train-kit

The learning half of an [ardegazu.ro](https://ardegazu.ro) bot — everything
game-agnostic between a gym and a live room. Live bots stay inference-only:
they replay plain-JSON MLP checkpoints with a 40-line forward pass (no ML
framework in the bot process) and append every decision to episode JSONL. A
separate trainer ingests those logs, runs offline self-play in
[ardegazu-peer-kit](https://git.ardegazu.ro/peer-kit/)'s gyms, fits with
tfjs-node, and ships a model only through a promotion gate — until it beats
the stock baseline, the bots keep playing the baseline.

- **`model`** — the checkpoint shapes: `MlpModel<G>` (a value/Q net as plain
  JSON) and `TabularModel<G>` (state-key → Q rows), structural guards, and
  ε-greedy `pickAction`.
- **`rev`** — `modelVersion(m)` → `"<episodes>@<rev8>"`, the string a bot
  advertises for the brain it plays; `"stock"` for anything it cannot label (no
  model, an invalid checkpoint, a nonsense episode count) — it is safe to call
  from a beacon loop. `modelRev(m)` → the 8 hex chars alone, and it throws on
  a non-checkpoint: the trainer/gate side wants the loud failure. The rev is sha256 over a
  deterministic serialization of the **weights only**: `updatedAt` moving does
  not move it, any weight change does, and two machines agree (MLP weights
  walked by index, tabular state keys sorted). Always computed from the
  weights — a `rev` stored in a checkpoint is greppable traceability, never
  trusted.
- **`forward`** — `mlpForward`/`mlpInit`: the whole live inference engine.
  Relu hidden layers, linear output — exactly what the trainer exports.
- **`episodes`** — append-only transition JSONL (`<game>-<YYYY-MM>.jsonl`),
  buffered per match, plus the byte-offset-resumable trainer-side reader.
- **`checkpoint`** — atomic checkpoint files (temp+rename; one writer, many
  readers, no locks) and `CheckpointStore`, the bots' interval reloader. The
  written JSON is `{v, game, episodes, rev, updatedAt, net|q}` in exactly that
  key order; `rev` (since v2.1.0) is additive and optional — files written
  before it load unchanged, and no reader requires it.
- **`tabular`** — Bellman updates over a string-encoded state space.
- **`dqn`** — the frozen-target pieces net games share: `netQs` per
  candidate, `dqnTargets`, a newest-kept `capReplay`.
- **`tf`** — MlpJson ↔ tf.Sequential and `fitNet`. `@tensorflow/tfjs-node`
  is an **optional peer dependency** reachable only through this module's
  dynamic import — inference-side consumers never load tensorflow.
- **`gate`** — the promotion file dance: `resolveModel` (resume from the
  promoted checkpoint, else the still-improving candidate) and
  `gatePromotion` (live file when promoted, `<game>.candidate.json` when
  not). What "beats the baseline" *means* stays with your game.
- **`ingest`** — per-bot episode-dir globbing and the persisted ingest
  offsets, so every live transition trains exactly once.

## Consuming

```jsonc
// package.json
"dependencies": { "ardegazu-train-kit": "git+https://git.ardegazu.ro/train-kit.git#v2.1.0" }
```

The package ships both the ClojureScript source (`src/ardegazu/train/`, for
CLJS classpaths) and committed deterministic ESM (`dist/`, what the `exports`
map serves, with hand-authored `types/index.d.ts`) — plain `node` works with
no build step. Zero runtime dependencies.

Inference side (what a live bot does):

```ts
import { CheckpointStore, checkpointPath, isMlpModel, modelVersion, netQs, pickAction } from "ardegazu-train-kit";

const store = new CheckpointStore(checkpointPath(modelDir, "valley-blocks"), 60_000);
// per decision: enumerate candidates, one feature vector each …
const m = store.model;
log(`brain ${modelVersion(m)}`); // "13208@a1b2c3d4", or "stock" while unpromoted
const a = m && isMlpModel(m, "valley-blocks")
  ? pickAction(netQs(m.net, feats), epsilon)
  : stockFallback(candidates); // the same baseline the gate measures
```

Trainer side (one process owns every weight update):

```ts
import { resolveModel, isMlpModel, mlpInit, dqnTargets, fitNet, capReplay, gatePromotion } from "ardegazu-train-kit";

const prev = resolveModel(modelDir, game);
const model = prev && isMlpModel(prev, game) ? prev : fresh(mlpInit([12, 32, 1]));
// … collect transitions from a peer-kit gym vs the stock baseline …
const { xs, ys } = dqnTargets(model.net, replay, gamma);
model.net = await fitNet(model.net, xs, ys, { lr: 0.001, epochs: 2 });
capReplay(replay, 50_000);
// … your eval vs the baseline decides `promoted` …
gatePromotion(modelDir, game, model, promoted);
```

`examples/train-gym.ts` is the whole loop in miniature — valley-blocks
self-play, three fit rounds, the gate line — runnable with the dev deps
installed. `docs/NEW-BRAIN-PROMPT.md` is the full recipe for teaching the
bot fleet a new game.

### Working on train-kit itself

```sh
npm install
npm run check    # shadow-cljs compile + tsc smoke of types/ (JVM ≥ 17 + clojure CLI needed)
npm run build    # → dist/ (committed)
npm test         # node --test; the tfjs tests skip where the native dep is absent
```

## Releasing

```sh
npm run check && npm run build && npm test
git commit … && git tag vX.Y.Z
cd ../git && ./assemble.sh train-kit
```

## License

MIT.

static mirror of HEAD · about · clone: git clone https://git.ardegazu.ro/train-kit.git