train-kit / docs / NEW-BRAIN-PROMPT.md
  1
  2
  3
  4
  5
  6
  7
  8
  9
 10
 11
 12
 13
 14
 15
 16
 17
 18
 19
 20
 21
 22
 23
 24
 25
 26
 27
 28
 29
 30
 31
 32
 33
 34
 35
 36
 37
 38
 39
 40
 41
 42
 43
 44
 45
 46
 47
 48
 49
 50
 51
 52
 53
 54
 55
 56
 57
 58
 59
 60
 61
 62
 63
 64
 65
 66
 67
 68
 69
 70
 71
 72
 73
 74
 75
 76
 77
 78
 79
 80
 81
 82
 83
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
# PROMPT: teach the ardegazu bot fleet a new game

You are an agentic developer working inside the ardegazu.ro workspace. The
bot fleet already learns séance, neon-grid and valley-blocks; your job is to
give it a real, learning brain for one more suite game. Everything you need
is in ardegazu-peer-kit (gyms, live clients, host sims) and this kit (the
learning core). Follow the flow in order. Verify each stage before moving on.

---

## 1. BRAIN SPEC (filled in by the requester)

```
GAME NAME:      <the peer-kit client name, e.g. "tessera">
DECISION POINT: <when the bot chooses: per tick / per piece / per turn>
CANDIDATES:     <the enumerable actions at a decision, and whether each has a
                 computable afterstate (board-after-action) — afterstates make
                 the strongest features>
REWARD:         <what shaped reward flows between decisions (score deltas,
                 survival ticks) and what the terminals pay (win/+, out/−)>
FALLBACK:       <the hand-written stock heuristic the net must beat — this is
                 BOTH what the bots play before promotion AND the gate baseline>
```

## 2. House architecture (non-negotiable)

- **One trainer process owns every weight update.** Live bots are
  inference-only: `mlpForward` over a plain-JSON checkpoint, ε-greedy
  `pickAction`, transitions appended to episode JSONL via `EpisodeWriter`.
  No tensorflow ever loads in a bot process — the kit's `tf` module is
  reachable only through a dynamic import the trainer alone takes.
- **Gyms live in peer-kit, views field-identical to the live clients.** A
  policy trained offline plugs into the live hook unchanged. If your game has
  no gym yet, that porting comes first (stage 3).
- **A model ships only through the promotion gate.** The trainer always
  resumes from the strongest thing it has (`resolveModel`: promoted else
  candidate), and writes back through `gatePromotion`. Until promotion, live
  bots play the stock fallback — the exact baseline the gate measures, so
  "promoted" means "measurably better than what's already live".
- **Checkpoints are shared files** (`<modelDir>/<game>.json`), reloaded by
  every bot within a minute (`CheckpointStore`). One writer, many readers,
  no locks. Log the brain a bot is playing as `modelVersion(m)` —
  `"<episodes>@<rev8>"`, or `"stock"` before the first promotion: the rev
  hashes the weights alone, so two bots quoting the same string really are
  running the same net, and a promotion visibly changes it.

## 3. Port the sim + gym into peer-kit (skip if they exist)

Every gym wraps a pure, step()-able, seedable sim ported from the browser
game — `ported-from:` provenance headers name the exact source file and sha.
Two shapes exist:

- **Host-authoritative games** (séance, neon-grid): the host sim owns all
  state; the gym steps it synchronously and re-folds the emitted frames into
  client-identical views. Reference: `peer-kit/src/ardegazu/peer/gym/neon_grid.cljs`.
- **Self-simulated games** (valley-blocks): each seat runs its own engine
  from a shared seed and the host only referees — the gym runs N engines in
  lockstep and feeds the real referee sim. Reference:
  `peer-kit/src/ardegazu/peer/gym/valley_blocks.cljs`.

Write determinism tests (same seed + same actions → identical episodes) —
everything downstream leans on them. Release peer-kit (`ardz kit-release
peer-kit`) before continuing.

## 4. Candidates + features in the bot

One normalized feature vector per candidate, trailing bias `1`, a
`<GAME>-FEATURES` constant. Compute features from **kit-shared metrics** (the
way valley-blocks uses peer-kit's `boardMetrics` on each candidate's
afterstate) so gym-trained weights transfer to live rooms byte-for-byte.
Reference implementation: `vb-features` in
`bot/src/bot/rl/valleyblocks_policy.cljs` — 12 features (`VB-FEATURES`) over
placement afterstates, net `[12, 32, 1]`.

Add the model guard to `bot/src/bot/rl/policy.cljs` — one line:
`(defn is-my-game-model [m] (isMlpModel m "my-game"))`.

## 5. The recorder (live inference + transition logging)

The pattern every recorder follows (reference: `make-recorder` in
`bot/src/bot/rl/valleyblocks_policy.cljs`):

- **On each decision hook call**: compute the candidate features; if a
  previous decision is pending, log its transition — reward is whatever
  accrued since that decision (score delta, survival ticks), `f2` = the new
  features, `done: false`. Then pick ε-greedy from the checkpoint (or the
  stock fallback when no model), remember `{feats, a}`, return the action.
- **On the game's terminal frame** (host `ov` or your own death): close the
  pending decision with the terminal reward (win bonus, top-out penalty),
  `f2: null, done: true`, and `flush()` the writer.
- **On round start**: reset the pending state.

Wire it through the policy registry next to the existing three: a
`bot/src/bot/policies/<game>.cljs` registering a provider (`:init` builds the
`CheckpointStore`, `:forRoom` returns `{:joinOptions {<hook>} :attach}`) —
`bot/src/bot/rooms.cljs` wires it in blind and calls `attach` after join. Add
the game to the `:games` list in `bot/src/bot/config.cljs` (the matchmaker,
`bot/src/bot/fleet/matchmaker.cljs`, cycles that list).

## 6. The trainer entry

Add a config block + DEFAULTS to `bot/src/bot/rl/train.cljs` and a
`train-my-game` mirroring `train-valley-blocks`:

1. Ingest live transitions (`readNewTransitions` — offsets make it
   exactly-once).
2. Gym self-play rounds with the current **frozen** net: seat 0 always
   learns ε-greedy; other seats alternate per episode between the current
   net and the stock fallback, so the learner feels real opposition instead
   of only its own reflection.
3. Fit: `dqnTargets(model.net, replay, gamma)` → `fitNet` → `capReplay`.
4. **Gate**: greedy net vs the stock fallback. Deterministic sims + two
   deterministic policies would replay ONE episode — add a dash of seeded
   jitter on both seats and alternate which seat the net takes. Promote at
   `winRate >= 0.5 && decided >= <floor>` via `gatePromotion`.

Expect a long candidate phase against a good hand-tuned fallback — that is
the gate doing its job. The hourly runs keep improving the candidate.

## 7. Deploy + verify

- Fleet: rsync the bot repo, `npm install` on the box, staggered
  `systemctl restart ardegazu-bot@<name>` (~20 s apart).
- Watch `journalctl -u 'ardegazu-bot@*'` for `joined <game> room … hosting`,
  a `selfplay <game> done winner=…` line, and `<game>-<YYYY-MM>.jsonl`
  growing under each bot's `episodes/`.
- Run the trainer once (`systemctl start ardegazu-trainer.service`) and
  expect the `<game>: … gate N% vs <fallback> → candidate` line plus
  `<game>.candidate.json` in the shared model dir.
- Promotion flips the live bots automatically on their next
  `CheckpointStore` reload — no restart needed.

static mirror of HEAD · about · clone: git clone https://git.ardegazu.ro/train-kit.git