1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112 | # ardegazu-train-kit
The learning half of an [ardegazu.ro](https://ardegazu.ro) bot — everything
game-agnostic between a gym and a live room. Live bots stay inference-only:
they replay plain-JSON MLP checkpoints with a 40-line forward pass (no ML
framework in the bot process) and append every decision to episode JSONL. A
separate trainer ingests those logs, runs offline self-play in
[ardegazu-peer-kit](https://git.ardegazu.ro/peer-kit/)'s gyms, fits with
tfjs-node, and ships a model only through a promotion gate — until it beats
the stock baseline, the bots keep playing the baseline.
- **`model`** — the checkpoint shapes: `MlpModel<G>` (a value/Q net as plain
JSON) and `TabularModel<G>` (state-key → Q rows), structural guards, and
ε-greedy `pickAction`.
- **`rev`** — `modelVersion(m)` → `"<episodes>@<rev8>"`, the string a bot
advertises for the brain it plays; `"stock"` for anything it cannot label (no
model, an invalid checkpoint, a nonsense episode count) — it is safe to call
from a beacon loop. `modelRev(m)` → the 8 hex chars alone, and it throws on
a non-checkpoint: the trainer/gate side wants the loud failure. The rev is sha256 over a
deterministic serialization of the **weights only**: `updatedAt` moving does
not move it, any weight change does, and two machines agree (MLP weights
walked by index, tabular state keys sorted). Always computed from the
weights — a `rev` stored in a checkpoint is greppable traceability, never
trusted.
- **`forward`** — `mlpForward`/`mlpInit`: the whole live inference engine.
Relu hidden layers, linear output — exactly what the trainer exports.
- **`episodes`** — append-only transition JSONL (`<game>-<YYYY-MM>.jsonl`),
buffered per match, plus the byte-offset-resumable trainer-side reader.
- **`checkpoint`** — atomic checkpoint files (temp+rename; one writer, many
readers, no locks) and `CheckpointStore`, the bots' interval reloader. The
written JSON is `{v, game, episodes, rev, updatedAt, net|q}` in exactly that
key order; `rev` (since v2.1.0) is additive and optional — files written
before it load unchanged, and no reader requires it.
- **`tabular`** — Bellman updates over a string-encoded state space.
- **`dqn`** — the frozen-target pieces net games share: `netQs` per
candidate, `dqnTargets`, a newest-kept `capReplay`.
- **`tf`** — MlpJson ↔ tf.Sequential and `fitNet`. `@tensorflow/tfjs-node`
is an **optional peer dependency** reachable only through this module's
dynamic import — inference-side consumers never load tensorflow.
- **`gate`** — the promotion file dance: `resolveModel` (resume from the
promoted checkpoint, else the still-improving candidate) and
`gatePromotion` (live file when promoted, `<game>.candidate.json` when
not). What "beats the baseline" *means* stays with your game.
- **`ingest`** — per-bot episode-dir globbing and the persisted ingest
offsets, so every live transition trains exactly once.
## Consuming
```jsonc
// package.json
"dependencies": { "ardegazu-train-kit": "git+https://git.ardegazu.ro/train-kit.git#v2.1.0" }
```
The package ships both the ClojureScript source (`src/ardegazu/train/`, for
CLJS classpaths) and committed deterministic ESM (`dist/`, what the `exports`
map serves, with hand-authored `types/index.d.ts`) — plain `node` works with
no build step. Zero runtime dependencies.
Inference side (what a live bot does):
```ts
import { CheckpointStore, checkpointPath, isMlpModel, modelVersion, netQs, pickAction } from "ardegazu-train-kit";
const store = new CheckpointStore(checkpointPath(modelDir, "valley-blocks"), 60_000);
// per decision: enumerate candidates, one feature vector each …
const m = store.model;
log(`brain ${modelVersion(m)}`); // "13208@a1b2c3d4", or "stock" while unpromoted
const a = m && isMlpModel(m, "valley-blocks")
? pickAction(netQs(m.net, feats), epsilon)
: stockFallback(candidates); // the same baseline the gate measures
```
Trainer side (one process owns every weight update):
```ts
import { resolveModel, isMlpModel, mlpInit, dqnTargets, fitNet, capReplay, gatePromotion } from "ardegazu-train-kit";
const prev = resolveModel(modelDir, game);
const model = prev && isMlpModel(prev, game) ? prev : fresh(mlpInit([12, 32, 1]));
// … collect transitions from a peer-kit gym vs the stock baseline …
const { xs, ys } = dqnTargets(model.net, replay, gamma);
model.net = await fitNet(model.net, xs, ys, { lr: 0.001, epochs: 2 });
capReplay(replay, 50_000);
// … your eval vs the baseline decides `promoted` …
gatePromotion(modelDir, game, model, promoted);
```
`examples/train-gym.ts` is the whole loop in miniature — valley-blocks
self-play, three fit rounds, the gate line — runnable with the dev deps
installed. `docs/NEW-BRAIN-PROMPT.md` is the full recipe for teaching the
bot fleet a new game.
### Working on train-kit itself
```sh
npm install
npm run check # shadow-cljs compile + tsc smoke of types/ (JVM ≥ 17 + clojure CLI needed)
npm run build # → dist/ (committed)
npm test # node --test; the tfjs tests skip where the native dep is absent
```
## Releasing
```sh
npm run check && npm run build && npm test
git commit … && git tag vX.Y.Z
cd ../git && ./assemble.sh train-kit
```
## License
MIT.
|