bot / README.md
  1
  2
  3
  4
  5
  6
  7
  8
  9
 10
 11
 12
 13
 14
 15
 16
 17
 18
 19
 20
 21
 22
 23
 24
 25
 26
 27
 28
 29
 30
 31
 32
 33
 34
 35
 36
 37
 38
 39
 40
 41
 42
 43
 44
 45
 46
 47
 48
 49
 50
 51
 52
 53
 54
 55
 56
 57
 58
 59
 60
 61
 62
 63
 64
 65
 66
 67
 68
 69
 70
 71
 72
 73
 74
 75
 76
 77
 78
 79
 80
 81
 82
 83
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
# ardegazu-bot

A resident of the [ardegazu.ro](https://ardegazu.ro) suite that lives on a
server: it plays the games, chats, draws on boards, and answers invites from
friends — a peer like any other, minus the browser. Run several and they
befriend each other, summon one another into games, and **learn to play** from
every match.

Built on the suite's Node kits:
[`ardegazu-peer-kit`](https://git.ardegazu.ro/peer-kit/) (games, identity,
social, hosting, the training gyms) and
[`ardegazu-rooms-kit`](https://git.ardegazu.ro/rooms-kit/) (chat + board).

## What it does

- Holds one **suite identity** (a seed file) and prints its **friend link**
  at startup — add it from the hub like any friend.
- **Presence**: it shows up as "here now" for its friends.
- **Invites summon it**: invite it to any suite app from the hub and it joins
  that room on the spot.
- **Games**: it plays all six — séance, neon-grid and valley-blocks with the learned
  policies below, the rest with their stock brains — and co-signs leaderboard
  receipts, so a match against it lands on the board.
- **Fleet**: given sibling friend links, bots befriend each other; the
  arranger (lowest pub) mints private rooms, invites partners, hosts the
  match itself (bot-only rooms actually run — peer-kit's ported host sims),
  and closes the room after one match.
- **Learning**: live bots are inference-only — they act from shared JSON
  checkpoints (a tiny built-in forward pass, no ML framework in the bot) and
  append every transition to episode logs. The **trainer** (hourly oneshot)
  ingests those logs, runs fast offline self-play in peer-kit's gyms (and,
  for the trading arm, a market gym over bursa's own pure fold), fits
  séance's tabular Q and the neon-grid/valley-blocks/bursa tfjs nets, and
  atomically rewrites the checkpoints; the bots reload them within a minute.
  The bursa gate is bespoke: a candidate promotes only by beating the
  promoted champion AND a random quoter on mean PnL over paired-seed
  episodes — the bootstrap spread quoter is the live fallback, never the
  baseline. The game-agnostic learning core (forward pass, episodes,
  checkpoints, fit helpers, gates) is
  [ardegazu-train-kit](https://git.ardegazu.ro/train-kit/).
- **Advertised brains**: over presence each bot names the brain it is playing
  per game — `brains: seance 59500@a625afba, valley-blocks stock` — using
  train-kit's `modelVersion`: episode count plus 8 hex chars of a sha256 over
  the checkpoint's WEIGHTS ONLY (so a metadata touch does not move it, two
  bots quoting one rev really run the same weights, and `stock` means "no
  promoted model, playing the fallback"). It rides one additive `br` key on
  the beacon, read from the live checkpoint store per beacon — a trainer
  promotion re-advertises itself within a reload window, no restart, no
  deploy, and logs one `brain <game> -> <version> (promoted)` line.
  *Not yet on the wire*: the bot hands the thunk to peer-kit's `Social`, and
  peer-kit's vendored copy of the social layer has to be refreshed to the
  version that carries `br` (and forward the option) before watchers see it —
  until then the log lines are the whole surface.
- **Chat**: replays history and echoes (swap the behavior for your own).
- **Board**: reads every element and signs the corner.

## Run

```sh
npm run stacks                    # the TWO runtime trees (see below)
# tfjs native addon (only the trainer needs it), once, if scripts are off:
(cd stack-a/node_modules/@tensorflow/tfjs-node && npx --yes node-pre-gyp install)
node stack-a/dist/main.js ./bot.json
```

**One compile, two npm roots.** The bot runs two generations of libp2p in one
process tree — peer-kit's v3 for the games, rooms-kit's v2 for the
chat/board/banca/bursa worker — and two builds of libdatachannel in one process
abort it. Node resolves a bare import from the importing FILE's directory
upward, so each stack gets its own package root and each bundle is emitted
inside it:

| root | libp2p | holds | bundles |
|---|---|---|---|
| `stack-a/` | v3 (peer-kit) | peer-kit, train-kit, tfjs-node | `dist/main.js`, `dist/rl/train.js`, `dist/fleet/link.js` |
| `stack-b/` | v2 (rooms-kit) | rooms-kit, id/social/wallet-kit, banca, bursa | `dist/worker.js` |

The repo root has no runtime tree at all — its only dependency is the compiler.
Compile time is still ONE project: one `deps.edn`, one `shadow-cljs.edn`, one
classpath; only the emitted files' locations and the npm trees differ.

Every kit — id, social, wallet, train, peer and rooms — is consumed as
**ClojureScript source** off one classpath, never as an npm dist. Only foreign
packages (libp2p, helia, orbit, tfjs, node builtins) stay bare imports, and
which generation those resolve to is decided entirely by which stack root the
bundle sits in.

The bot is ClojureScript (v2): both `dist/` directories are committed and are
what runs — the server never builds. Rebuilding (`npm run build`, four
shadow-cljs release builds) needs a JVM ≥ 17 and the `clojure` CLI on the dev
machine; `deploy/check-dist.sh` gates that the committed dists match a fresh
build, that no bundle carries the other stack's namespaces or names a package
that exists only in the other generation, that no kit is imported as a dist,
that each stack's tree holds exactly one libp2p at its own generation and not
the other's WebRTC native, and that every bare import each bundle emits
resolves INSIDE its own root.

`bot.json` (see `deploy/bot.example.json`): `rooms` pins permanent rooms
(usually empty — invites do the work), `fleet.siblings` lists the other bots'
friend-link fragments, `fleet.matchmaker.enabled` turns on self-play,
`learning.modelDir` points at the shared checkpoint directory.

With npm's `ignore-scripts` on — or after a flaky install — fetch the
chat/board native prebuild once:

```sh
(cd stack-b/node_modules/@ipshipyard/node-datachannel && npx prebuild-install -r napi)
```

## Deploy (systemd fleet)

One template unit runs any number of bots; a oneshot + timer runs the trainer.
All instances share a static user so the checkpoint dir is writable by the
trainer and readable by everyone.

```sh
rsync -a . server:/opt/ardegazu-bot/
ssh server 'cd /opt/ardegazu-bot &&
            sh deploy/install-fleet.sh duhul moroi strigoi iele zmeu'
```

`install-fleet.sh` now runs the two `npm ci`s itself (one per stack root), so
there is no separate install step and no way to end up with one tree fresh and
the other stale.

`install-fleet.sh` creates the `ardegazu-bot` user, keeps any existing seed
(migrating a pre-fleet flat state dir), mints the missing seeds, cross-links
every bot's friend link into its siblings' configs, installs
`ardegazu-bot@.service` + `ardegazu-trainer@.{service,timer}`, and starts the
fleet staggered. Logs: `journalctl -u ardegazu-bot@duhul -f`.

**Every bot has its own brain.** `modelDir` is per bot (`<state>/<bot>/models`)
and each gets its own `ardegazu-trainer@<bot>` instance reading only its own
episodes, so the brains diverge from a common ancestor rather than sharing one
checkpoint set — visible on the hub's `brains` line and in the leaderboards.
Gym self-play is the dominant data source and each instance runs its own gym at
full rate, so N brains do not learn N times slower; only the live-match ingest
is partitioned.

**A fleet can span hosts.** The args are the bots that run *here*; siblings
elsewhere come in through a links file, since their seeds live on the other
machine:

```sh
FLEET_EXTRA_LINKS=/path/to/links sh deploy/install-fleet.sh iele zmeu balaur
```

one `<name>=<friend-link>` per line. Without it each half would only see itself
and cross-host self-play would never be arranged.

Mind the managed relay's budget: ~20 connections per IP, one per social agent
plus one per open room. The matchmaker's `maxConcurrentSelfPlay` and the
room reaper (sibling rooms close after their match, human rooms after a long
idle) keep the fleet inside it.

Chat/board rooms run in per-room worker processes: the two kits' WebRTC
transports load two different builds of the same native library, which cannot
share a process — and a native crash restarts one room, not the bot.

### Watching a room without touching it

```sh
node scripts/floor-observer.mjs                          # the bursa floor
node scripts/floor-observer.mjs <secret> <salt> <origin> [window-sec]
```

A read-only diagnostic that joins a rooms-app room on the bots' own stack with
a throwaway in-memory identity, replays the full history and reports peers,
hellos and entries for a two-minute window (`REPORT …` lines at the end). It
never appends — no hello, no publish, no deposit — so it is safe to point at
any production room when a market looks stale or a bot looks deaf; state goes
to a throwaway tmpdir.

### Backup and restore

```sh
sh deploy/backup-fleet.sh <host> <out-dir>                  # ~2 MB
sh deploy/backup-fleet.sh <host> <out-dir> --with-episodes  # ~1.4 GB
sh deploy/restore-fleet.sh <archive> <host>                 # plan only
sh deploy/restore-fleet.sh <archive> <host> --yes           # apply
```

One tarball that can rebuild the fleet on another VPS or roll it back. It is
streamed over ssh, so nothing is left on the server, and it lands mode 0600 —
it holds the identity seeds.

The state dir is ~1.4 GB and ~99% of that is `episodes/`, the training JSONL
the fleet regenerates by playing, so it is opt-in. The default carries what
you cannot get back any other way: the seeds (44 bytes each, and the seed IS
the bot's suite identity — lose it and every friend link it handed out is
dead), `kv.json` (friends, blocks, inbox cursors), `rooms/`, the trained
`models/`, the per-bot configs, the unit files, and `package*.json` so a new
host resolves the same sha-pinned kits. A MANIFEST records the seed checksums,
so a restore can be verified without ever printing a seed.

`restore-fleet.sh` **plans by default and writes only with `--yes`**, because a
wrong restore silently replaces identities. It reports which host seeds differ
from the archive's before touching anything. `--keep-seeds` restores everything
except the seeds; `--no-start` leaves the units stopped. The code is not in the
archive — deploy it first (rsync + `install-fleet.sh`, above), since the VPS
never builds ClojureScript.

## Training loop

```
bots (inference + episode JSONL)  →  trainer (hourly: ingest + gym self-play)
        ↑                                        ↓
        └──────── shared models/*.json ←─────────┘
```

Run the trainer by hand any time: `node stack-a/dist/rl/train.js --config
/etc/ardegazu-bot/trainer.json [--game seance|neon-grid|valley-blocks|bursa]
[--gym-only]`.
Checkpoints are plain JSON and safe to inspect; delete one to fall back to
the stock brain until the next training run.

Public mirror (git dumb-HTTP — no git server):

```sh
git clone https://git.ardegazu.ro/bot.git
```

static mirror of HEAD · about · clone: git clone https://git.ardegazu.ro/bot.git