Headless mode (server)
Note
This section describes an advanced, optional mode of blunderDB, intended for server deployments, multi-user setups and automation. The normal and recommended use of blunderDB remains the desktop application described in the previous chapters. If you use blunderDB on your own, on your computer, you do not need this mode: you can skip this chapter without losing any of the analysis features.
Overview
In addition to the desktop application and the command-line commands (see Command Line Interface (CLI)), the same blunderdb binary can run in headless mode: without a graphical interface, driven entirely from the command line or over the network. This mode covers three use cases:
the
servedaemon — exposes the blunderDB engine as an HTTP + JSON service, to run a shared database on a server and access it from several clients;the generic
calldispatcher — calls any storage operation directly, locally, for scripting and testing;the
migratecommand — transfers a single-user SQLite database to a multi-user PostgreSQL backend.
These three use cases rely on a shared storage layer that can talk to two backends: SQLite (the usual .db file format of the desktop application) and PostgreSQL (for multi-user server deployments).
The serve daemon
blunderdb serve runs the engine as an HTTP service that responds in JSON. It lets you host a position database on one machine and access it from several clients.
# sqlite
blunderdb serve --db database.db --addr 127.0.0.1:8080
# postgres
blunderdb serve --backend postgres \
--dsn "postgres://user:pass@host:5432/blunderdb?sslmode=disable" \
--addr 127.0.0.1:8080
Note
sslmode=disable is only suitable for a trusted private network — a database in a neighbouring container, on a network with no route to either the host or the Internet. For a remote database, sslmode=require encrypts the link, and verify-full also checks the server’s certificate and hostname. The other connection strings on this page carry sslmode=disable for the same reason: they all describe a private network.
Warning
The daemon performs no authentication. It trusts the X-Tenant-ID request header and must run behind a reverse proxy (nginx, Caddy…) responsible for authentication. Never expose it directly to the public Internet.
X-Tenant-ID is the tenant’s integer (1, 2, 42…): it is the reverse-proxy’s job to map the authenticated account to that integer. A name (alice) is refused with 400 invalid, never converted.
Options:
Option |
Default |
Meaning |
|---|---|---|
|
– |
SQLite file (shorthand for |
|
|
storage backend: |
|
|
backend connection string |
|
|
listen address |
|
|
log level: |
|
|
exposes |
|
|
serves the read-mostly web page under |
|
|
serves the tournament and event direction gestures; off by default, see Direction gestures |
|
|
offers the write tools of |
|
|
serves the transcription moves ( |
|
|
closes a transcription session that has been idle for longer than this |
|
– |
enable CORS for this origin, a comma-separated list of origins, or |
|
|
requests per second per tenant limit (0 = disabled); enabled by default at a generous value rather than opt-in, so a compose file that only thinks about the database does not end up inheriting a daemon with no limit at all |
|
|
token-bucket size for request bursts |
|
|
PostgreSQL: enables per-tenant Row-Level Security (defence in depth, opt-in) |
|
– |
optional two-sided bearoff database ( |
|
– |
directory for the daemon’s signing identity (created on first use); needed for |
|
– |
serves the |
|
– |
exposes |
Most options can also be given through an environment variable (BLUNDERDB_BACKEND, BLUNDERDB_DSN, BLUNDERDB_ADDR, BLUNDERDB_LOG_LEVEL, BLUNDERDB_METRICS, BLUNDERDB_CORS_ALLOW_ORIGIN, BLUNDERDB_RATE_LIMIT_RPS, BLUNDERDB_RATE_LIMIT_BURST, BLUNDERDB_RLS, BLUNDERDB_TS_PATH, BLUNDERDB_IDENTITY_DIR, BLUNDERDB_OPS_ADDR, BLUNDERDB_PPROF_ADDR): an explicit flag still wins over the matching variable.
The daemon has no data-directory option: it writes its bearoff tables to $XDG_DATA_HOME/blunderdb, or ~/.local/share/blunderdb otherwise. So it is XDG_DATA_HOME that moves them — see The bearoff databases.
The rate limiter’s own bucket table carries a hard cap (10,000 distinct tenants): beyond it, each new tenant evicts the least-recently-used bucket rather than letting the table grow without bound — useful if a client sends many distinct X-Tenant-ID values, whether by accident or not, between two periodic sweeps of idle buckets.
blunderdb serve now rejects any unexpected positional argument (beyond the single leading serve that an ENTRYPOINT already reduced to the bare binary lets through): without this check, a flag placed after such an argument was silently ignored — docker run image serve --addr :9090, a natural reflex since the image’s ENTRYPOINT is already serve, used to start on :8080 without a word.
Endpoints
The service exposes operational endpoints, always present:
GET /healthz— liveness (the process is running);GET /readyz— readiness (storage responds and its schema is at the expected version);GET /metrics— Prometheus metrics (if--metricsis enabled);GET /app/— the read-mostly web page (if--webis enabled).
The web page
blunderdb serve --web serves a page under /app/: a library you can consult from a tablet or a phone, with nothing to install.
It can do three things, and that list is the decision, not a stage:
consult a position, its analysis and its board;
search, with the same token grammar as the application’s command line;
review an Anki deck — answer revealed and grade given.
It cannot edit a position, import, delete, manage collections, matches, tournaments or the configuration, and it will not learn to. A feature missing here is not a gap: it is the perimeter.
It is off by default, and that default is the decision. The daemon authenticates nobody: it trusts the X-Tenant-ID header and must run behind a proxy that authenticates. Shipping a browser-reachable interface switched on out of the box would invite exactly the deployment that rule forbids.
The page sends no tenant: the proxy sets the header, as for any other client. In local development, and there only, /app/?tenant=1 names one — which changes nothing about the safety of a daemon that already accepts that header from anyone.
The page’s files are served without a tenant, deliberately: a browser must be able to load the page before the proxy assigns it anything, and a page holds no data.
Liveness and readiness answer two different questions. /healthz always answers 200 as soon as the process serves requests, without ever querying the storage: an orchestrator restarts a container whose liveness fails, and a briefly unreachable database must not restart a healthy daemon in a loop. /readyz answers 503 (with status set to down or version_mismatch) as long as the database does not answer or its schema is not the binary’s: traffic is simply routed away until it comes back.
The blunderdb healthcheck subcommand (also present in the serve binary of the container image) performs a GET /readyz request on the local daemon and returns 0 when it is ready, 1 otherwise; the address is that of --addr or BLUNDERDB_ADDR, :8080 by default. It is the Docker image’s HEALTHCHECK, and it works just as well from a script or a systemd unit:
blunderdb healthcheck --addr 127.0.0.1:8080 && echo ready
The business surface follows the POST /v1/<family>.<method> shape (for example /v1/positions.save, /v1/matches.get). The families cover positions, analyses, matches, comments, collections, tournaments, Anki cards, filters, sessions, history (search and commands), search, metadata, library settings, statistics, import and export. Listing endpoints return an NDJSON stream (one JSON object per line). The server shuts down cleanly on SIGINT / SIGTERM.
An error returns the envelope {"error":{"code":…,"message":…}}. The code not_found says a named resource does not exist; unknown_route, also a 404, says the daemon does not serve the method called: a client and a daemon of different versions, or a family the daemon only serves behind a flag. A client concludes that data is missing only on not_found.
positions.save returns {"id":…,"created":…}. created is true for the one call that inserted the position, and the write itself says so: a client that copies a position then its analysis, and must undo the copy after a failure, deletes the position only if it created it, without the race of a prior positions.exists.
What /v1 promises
A client written against /v1 must keep working. The rule fits in three lines, and it is more useful written down than guessed:
What exists does not change meaning. A
/v1route is never renamed, removed or re-purposed. A request or response field is never renamed, dropped or retyped.What is added is added. A new route, an optional request field, a new field in a response: a client that ignores them keeps working — that is the definition of “compatible” taken here. A client must therefore ignore the fields it does not know rather than refuse them.
Everything else is
/v2. Making a field mandatory that was not, changing a unit, changing the meaning of an error code: those are breaks, and they live under another prefix, beside/v1, for as long as clients need to cross over.
Two clarifications that matter. The /ops/ routes are not covered: they operate a deployment, change with it, and are not an API for third-party programs. And the contract itself is generated from the daemon’s route table (openapi.yaml, API contract): it cannot describe anything other than what the server serves.
Transcribing through the API
The transcriptions.* family lets an external client transcribe a match move by move, with the same logic as the desktop. The reads (list, get, exportMat, losses) are always served. The moves (create, open, editMatch, apply, undo, redo, close, finish, abandon) are served only with serve --transcription: without this flag, these routes answer 404.
create and open return the draft’s state, its revision and a sessionId. apply, undo, redo, close and finish name that sessionId: missing → 400, expired or unknown session → 410; the client then reopens the draft (open), cursor at the end of the document. abandon names no session: it deletes the draft under the If-Match revision alone. Every move that writes carries the last revision seen in the If-Match header and returns the next one:
If-Matchmissing → 428;stale revision → 409; the error envelope gives the current revision (
details.revision) and the draft’s fresh state (details.state: document, revision, session and cursor), which the client displays before replaying its move if it still applies.
The revision advances only when the document changes (header and actions): moving the cursor or entering a die of the action in progress writes nothing and returns the same revision. A session belongs to the draft, not to a client: open returns the live session when there is one, and the tabs or workstations that share it also share the cursor and the undo stack.
The session keeps only the undo stack, the cursor and the entry in progress: the draft is written after every move that changes it, so a lost session (inactivity, restart, another instance) loses no move. transcriptions.get returns the revision as an ETag and answers 304 to an If-None-Match that names it.
finish saves the Match and deletes the draft, abandon deletes it without a Match, close only releases the session. editMatch opens a draft on an existing Match and, for an imported match, returns the count of analyses and comments that the transcription does not keep (losses.lossy). The analysis of the saved match is started with gammonnet.analyzeMissing.
Warning
The daemon authenticates nobody: opening writes means entrusting them to the proxy (Deployment behind an authenticating proxy). A “transcriber” role is a proxy rule on the /v1/transcriptions. prefix, not a notion of the daemon.
A Python client
clients/python/ holds a minimal client with no dependency beyond the standard library — the daemon speaks POST and JSON, which urllib and json cover entirely:
from blunderdb import Client
api = Client("http://127.0.0.1:8080", tenant=1)
print(api.metadata_counts())
for position in api.positions_list({"limit": 10}):
print(position["id"])
It comes in two halves, on purpose. _generated.py carries one method per route, generated from the daemon’s route table by go run ./cmd/openapi-gen: a hand-written surface would drift the day a route is added, and nobody would notice before a user did. client.py carries the transport — the session, the tenant header, the error envelope, the NDJSON decode — and is written by hand. What changes with the API is generated; what changes with judgement is not.
Method names are family_operation in snake_case: /v1/positions.loadByIds becomes positions_load_by_ids(). The family is kept because several families share an operation name (list, delete), and a bare list() would collide.
events() follows /v1/events and returns one dictionary per message (see Being notified of gestures: /v1/events).
A failure raises APIError, carrying the daemon’s envelope as it is: the code (what a program branches on), the message (what a person reads), the HTTP status and the details.
Embedding the engine in a Go program
pkg/blunderdb/server.Bootstrap opens the storage and returns a set of handlers in the calling process, without listening on a port. It is the way in for a trusted parent — gammonGo — that wants the position library without running a daemon beside it or speaking HTTP to itself.
What that assumes is explicit: the parent is trusted. There is no tenant to verify, no header to validate, no rate limiter — those belong to the daemon because it faces a network, and ADR-0005 says why. A program that embeds the engine chooses its own tenant and answers for its calls.
Tournament direction and events
Tournaments directed at the workstation and the events that group them (rencontre in the API and its /v1/rencontres.* routes) are read through the API, under the caller’s tenant, with the same code as the workstation. Reading is always served; the gestures (entering a result, pairing, creating an event) are served only under serve --direction (Direction gestures).
directions.listanddirections.directoryread the whole tenant: the list of directed tournaments, the player directory.The other
directions.*take{"tournamentId": N}:directions.get(the full view: proposals, standings, matches in progress),directions.participants,directions.freeParticipants,directions.tableGrid,directions.brackets,directions.standings,directions.standingsCsv,directions.history(optionalplayerandmatchfilters),directions.clock,directions.slots,directions.lastDecision,directions.pageHtmlanddirections.pairingSheetHtml(withround).rencontres.list, thenrencontres.getandrencontres.pageHtmlwith{"id": N}.rencontres.pageHtmlrenders the room’s wall page, a self-contained HTML document in thehtmlfield: a wall screen displays it and reloads it periodically.
Pages are rendered in French, the language of the direction engine. A tournament that is not directed, or that belongs to another tenant, answers 404.
Conditional reads. Each of these routes returns an ETag header. Sent back in If-None-Match, it gets 304 with no body as long as nothing the route reads has changed. Any write changes the ETag at once: a gesture in the tournament or in a tournament of the same event, the attachment of a match, a draft started from a slot, the renaming of a tournament, a change to the event. Answering 304 replays no tournament, which makes a wall page that polls every few seconds inexpensive. Only what depends on the time is an exception: proposals, the clock and pages are computed at the moment of the read, so an ETag is valid for one minute at most. A client that reloads thus sees a deadline or a pause go by within the minute.
These routes are POST. For this verb, RFC 9110 (§13.1.2) answers 412 to a matching If-None-Match. The daemon answers 304 nonetheless: the request body carries only the parameters of a read with no effect, which behaves like a GET. The If-None-Match: * form is refused (400), because it designates no response the client would already have. An invalid request (a negative round, for example) is refused before any condition.
curl -si -X POST http://127.0.0.1:8080/v1/rencontres.pageHtml \
-H 'X-Tenant-ID: 1' -d '{"id":1}' | grep -i '^etag'
curl -si -X POST http://127.0.0.1:8080/v1/rencontres.pageHtml \
-H 'X-Tenant-ID: 1' -H 'If-None-Match: W/"…"' -d '{"id":1}'
# HTTP/1.1 304 Not Modified
Like the rest of /v1, these routes authenticate no one: behind the proxy (Deployment behind an authenticating proxy), anyone who reaches a tenant’s /v1/directions. prefix reads its tournaments, player names included. A proxy that reserves these reads for certain users does so with a rule on that prefix and on /v1/rencontres..
Direction gestures
blunderdb serve --direction opens the gestures the workstation performs on a directed tournament and on an event. Without this flag, these routes answer 404, as if absent. call always serves them.
directions.create(tournamentId,config,seed),directions.setConfiganddirections.previewConfig(config, the configuration in the engine’s JSON format);registrations:
directions.enterParticipants(players),directions.addParticipant(name,club,rating; withsectionandkey, a latecomer takes a bye slot),directions.updateParticipant,directions.withdraw,directions.reinstate,directions.makeAbsent,directions.makeAvailable,directions.addPair,directions.updatePair;the course of play:
directions.confirmProposal(action, as proposed bydirections.get),directions.confirmAllProposals,directions.startMatch,directions.enterResult,directions.enterForfeit,directions.moveMatchToTable,directions.cancelMatch,directions.correctResult,directions.close,directions.reopen,directions.addNote,directions.attachMatch,directions.detachMatch;the event:
rencontres.create,rencontres.update,rencontres.attach,rencontres.detach,rencontres.trash,rencontres.setTableOutOfService,rencontres.setBreaks;the table properties:
rencontres.setTables(id,tableSettings, one entry per table that carries any: number, name, room, reserved, assigned to),rencontres.setEventRooms(id,tournamentId,rooms, the rooms the competition plays in; none means all tables) anddirections.setTables(tournamentId,tableSettings) for a competition that plays alone.
A tournament gesture returns the full view of the tournament, like directions.get; an event gesture returns the event. The service then rewrites the display pages in the folder the database designates, as at the workstation. A page that cannot be written (folder gone, disk full) does not cancel the gesture: the response carries a Direction-Page-Warning header for each page not written (tournament 3, rencontre 2), without the server path, and the workstation shows it in its status bar.
A gesture that the rules refuse (empty name, occupied table, tournament that has not started, configuration rejected by the engine) returns 400 with the reason. A failure of the daemon or its database returns 500, without detail: the reason stays in the daemon’s log.
Version required. Every read of a tournament or an event returns a Direction-Version header, and every gesture sends it back in If-Match:
without
If-Match(or with*), the gesture is refused:428;if someone has written since that read, the gesture is refused:
409. The error’sdetailsfield carries the fresh state and itsversion: the client re-reads, then replays its gesture if it is still valid;otherwise the gesture is applied and returns the new version in
Direction-Version.
The comparison is made in the gesture’s transaction, under a database lock (PostgreSQL advisory lock per tournament or per event, SQLite write lock): of two gestures sent on the same read, only one is applied, whether they go through the same daemon, through two daemons on the same PostgreSQL database, or through the workstation and call on the same file. The gesture is written all or nothing. A tournament played in an event has its event’s version, so a gesture in a sibling competition changes it too. directions.create and rencontres.create target nothing existing and take no version.
Idempotence. A gesture that carries an Idempotency-Key header is applied only once: sent again with the same key, it returns the first response, with its headers (Direction-Version included) and Idempotency-Replayed: true. A double click or a network retry does not enter two results; two simultaneous sends of the same key run the gesture only once. Only a successful response is kept.
The key is bound to the request body: the same key with another body returns
422.The replay comes before the version check: it returns the kept response without
428or409, even if the version has moved since.Keys live in memory, in each instance of the daemon, for 24 hours, at most 1,000 per tenant: a restart forgets them, and another instance does not know them.
curl -si -X POST http://127.0.0.1:8080/v1/directions.get \
-H 'X-Tenant-ID: 1' -d '{"tournamentId":3}' | grep -i '^direction-version'
curl -s -X POST http://127.0.0.1:8080/v1/directions.enterResult \
-H 'X-Tenant-ID: 1' -H 'If-Match: "…"' -H 'Idempotency-Key: t4-r2' \
-d '{"tournamentId":3,"matchId":"m7","winner":"aa","scoreA":7,"scoreB":3}'
Warning
The daemon authenticates no one (ADR-0005). With --direction, anyone the proxy lets through enters results. The engine knows no role (director, referee, reader): a role is a proxy rule, which reserves /v1/directions. and /v1/rencontres. for directors, or lets only reads through. Never run --direction on a reachable daemon without that proxy, even on a club’s Wi-Fi.
Being notified of gestures: /v1/events
GET /v1/events is a Server-Sent Events stream (text/event-stream): one message per validated gesture of the tenant, published after the database write, never for a refused or cancelled gesture. The message says what changed and its new version, not the state: the client re-reads what it displays, with If-None-Match.
event: rencontre—rencontreId,tournamentIds(the event’s competitions, before and after the gesture) andversion;event: direction—tournamentIdandversion, for a tournament played outside any event;event: transcription—transcriptionIdandrevision; an abandoned or finished draft carriesremoved(andmatchIdfor Finish).
removed: true signals what no longer exists. The route is served only with --direction or --transcription: without them, the daemon writes nothing it would have to announce, and /v1/events answers 404. Like every /v1/ route, it requires X-Tenant-ID: a subscriber hears only its own tenant. A tenant holds at most 16 open streams at a time; beyond that, 429. The workstation uses the same service but plugs no bus into it: its gestures are not announced.
The tournament, rencontre and transcription parameters (comma-separated or repeated identifiers) narrow the subscription: a message passes if it names one of them. A tournament of an event receives the messages of its event. An unknown parameter or an invalid identifier returns 400.
curl -N http://127.0.0.1:8080/v1/events?rencontre=2 -H 'X-Tenant-ID: 1'
No history. The daemon keeps no message. Every stream opens with event: resync, carrying an id: the client may have missed gestures before connecting, or between two connections, and re-reads everything it displays. The reason is reconnected when the request carries Last-Event-ID, subscribed otherwise. A subscriber that is too slow, whose queue of 64 messages is full, is disconnected after the same resync: it never delays a gesture. The stream announces a reconnection delay of 3 seconds.
Through a proxy. A : ping comment is sent every 25 seconds so that a proxy does not cut a silent stream; X-Accel-Buffering: no asks nginx not to buffer it. The stream is not compressed, escapes the timeout of ordinary requests and counts as only one request for rate limiting. Stopping the daemon closes all streams; a subscription requested during shutdown receives 503.
Multiple instances. On SQLite, a single instance holds the database: the in-memory bus is enough. On PostgreSQL, as soon as --direction or --transcription is active, each instance relays its actions to the others through LISTEN/NOTIFY, on the blunderdb_events channel: a subscriber attached to one instance hears an action validated on another, or issued through call on the same database. The tenant travels in the notification and the instance that receives it delivers it only to that tenant’s subscribers. Each instance opens two more connections (application_name blunderdb-events-… for listening, blunderdb-notify-… for sending); an instance that cannot listen at startup refuses to start. call announces without listening, and serves its request even if it cannot announce.
Any role allowed to connect can emit on this channel, under --rls too. A received notification is trusted only if its tenant is valid and its kind known; the rest is logged and ignored. At worst, a forged notification can make a tenant’s subscribers re-read their data.
The notification is sent after the database write, like the local message. Two losses remain without
resync: an instance killed between the write and the notification, and a shutdown that cannot send in 2 seconds what is left in the queue. The action is validated, but the streams already open on the other instances only learn of it when their client reconnects.A lost listening connection is re-established, with a growing delay from 250 ms to 30 s. Actions from other instances that happened during the outage are lost: on recovery, each subscriber of the instance receives a
resyncwith the reasonmissed. A notification too long for PostgreSQL (8,000 bytes), or that an instance could not send, reaches the others as this sameresyncfor the tenant concerned.The stream
idvalues are specific to each instance. A client that a load balancer sends to another instance gains nothing from them: theresyncthat opens every stream makes it re-read what it displays.
The bearoff databases
The daemon computes its two default tables at start-up, in the background (TS-06-06 for the cube verdict, OS-06 for the EPC): about six seconds of one core, once, in its data directory — $XDG_DATA_HOME/blunderdb, or ~/.local/share/blunderdb otherwise. Nothing is downloaded and nothing is embedded in the binary (ADR-0027). If that directory is read-only, the tables are held in memory for the life of the process: the service starts, it simply pays for the computation at every restart.
A wider domain is not computed at start-up — TS-06-11 weighs 1.2 GB and takes minutes, which is not something a service decides on its own. It is up to the operator to make it, with the CLI, in the volume the daemon will read:
# generate
blunderdb bearoff generate --ts 6x11 --data-dir /srv/data/blunderdb
# serve
XDG_DATA_HOME=/srv/data blunderdb serve --db database.db
blunderdb serve --db database.db \
--bearoff-ts /srv/data/blunderdb/gnubg_ts6x11.bd
The first launch lets the daemon find the table on its own in its data directory; the second points to it by path, wherever it is. --data-dir is an option of the bearoff subcommands, never of serve.
blunderdb bearoff list --data-dir /srv/data/blunderdb says what the volume holds and what each domain would cost; blunderdb bearoff verify exits with an error on a corrupt table, which makes it a start-up probe usable as is. See Command Line Interface (CLI) for the details.
The operator routes
Two calls do not stop at the tenant making them, and therefore live under a prefix of their own, POST /ops/<family>.<method>:
/ops/maintenance.vacuum(SQLite backend) rewrites the whole file, every tenant’s data included, and holds a write lock for the duration;/ops/tenant.purge(PostgreSQL backend) destroys a tenant’s data, and the tenant it destroys is the one named in the header the caller controls.
The daemon authenticates nobody (see below): a route reachable by one tenant is a route every tenant may call. The prefix exists so the proxy can refuse both with a single rule. Never expose /ops/ through the public proxy. Under nginx, the rule fits in one line of the server block; under Caddy, in two lines of the site:
location /ops/ { return 403; }
location /metrics { return 403; }
@closed path /ops/* /metrics
respond @closed 403
The --ops-addr <host:port> option goes further: the two routes then leave the --addr address and are served only on that second listener, to be bound on an administration interface. Without the option they stay on the main listener, and blocking them is the proxy’s business.
These routes require the X-Tenant-ID header like every other one — a purge names the tenant it destroys, and needs that header more than anything else. Only the probes (/healthz, /readyz) and /metrics go without.
This is why the denial rule above also covers /metrics: requiring no tenant, it is readable by anyone who reaches the daemon, and it publishes the database size and the work in flight, across every tenant. It is meant to be consulted from the daemon’s own machine, or through a path the proxy reserves for operations. The third thing that must never be exposed is not a route but a listener: the one for --pprof-addr, which has no notion of tenant and hands out a profile of the whole process. It should bind to an administration interface, never published by the proxy.
What did not move under /ops/: /v1/gammonnet.sweepStale. The catch-up is expensive but it is scoped to the calling tenant; what bounds it is the rate limit and the in-flight gauges, not a trust boundary.
The complete contract — every method, its request and its response — is generated from the source and versioned: openapi.yaml at the repository root (OpenAPI format, schemas included) and its readable annex, API contract (one table per family). Both are regenerated by go run ./cmd/openapi-gen, and a dedicated test fails if either falls behind the routes actually registered.
Every /v1 request accepts a JSON body (Content-Type: application/json, or no header at all — a body of another type is refused with 400 invalid rather than failing on a confusing JSON parse error); a known method called with the wrong HTTP verb answers 405, the Allow header naming the one verb it accepts. Listing methods that take a limit refuse anything past 1000 rows per page (400 invalid) rather than honouring an unbounded value.
Every listing family accepts limit and offset: positions.list, positions.listIds, matches.list, search.find, anki.reviewLog, comments.listAll, tournaments.list and collections.positions. Both default to zero, which means what it has always meant: everything. There is no implicit cap — a stream is not held in memory, so an unbounded list costs time and bandwidth but never the daemon’s footing, whereas a silent default limit would have a client read a truncated list believing it complete. What the two parameters buy is the ability to page, for whoever wants to.
Every TCP connection is bounded in read/write time per request — a generous budget for ordinary calls, a far larger one for streaming routes (NDJSON lists, imports/exports, the gammonNet catch-up sweep) — and the number of them open at once is capped: past that many, a further connection waits for one of the existing ones to free up rather than every connection unconditionally getting its own thread of execution. A graceful shutdown (SIGINT/SIGTERM) first cancels every in-flight import and gammonNet catch-up — each answers with a trailing {"event":"cancelled"} event instead of having its connection cut with no explanation — before closing the server within the usual grace period. An uploaded import’s temporary file keeps only the extensions the daemon recognises from the original one (.xg, .xgp, .sgf, .mat, .bgf, .ogxm, .txt, .db, .dbx), and every import in flight — across every tenant — shares one global quota of bytes spooled to disk: past it, a new import is refused (too many requests) rather than letting $TMPDIR usage grow without bound.
/v1/imports.json reads back a blunderDB JSON export by filling gaps: the analysis it carries is written only on a position that has none yet, never replacing an existing analysis, and the rollouts on both sides are kept.
The search family offers three doors onto the same search. search.find takes the complete filter object, field by field. search.query takes a query written in the language of the application’s command bar (s cube p>30 E>50, described in List of commands) and streams the same positions; it is the only way to reach, over the network, the filters that have no obvious field — move pattern, comment text, player, date, excluded dice, zones and blots. search.parse searches nothing: it answers what a query means — the filters it denotes, its canonical form (two equivalent queries share it, which makes a saved search comparable) and its diagnostics.
A query carrying a token nothing recognises is refused (400 invalid, naming the token) rather than run while narrowing the search in silence. A token that is understood but has no effect here — x, which switches on the exclusion structure, itself a board rather than text — travels in the X-BlunderDB-Query-Diagnostics header, so that the body stays NDJSON positions for every existing client.
Two methods in the positions family decode a position without storing it: positions.fromXGID rebuilds a position from an XGID string, and positions.fromXGP from a single-position .xgp file.
POST /v1/exports.sqlite exports the entire current tenant — positions, collections, matches, tournaments, analyses, comments, played moves, filter library and Anki decks — into an SQLite file that can be opened as-is on the desktop; the export always carries the whole database, selecting a subset is a desktop or CLI action. The request’s JSON body is optional and only accepts watermarkOrigin / watermarkNote, to apply a watermark signed with the daemon’s own identity (--identity-dir) — without these fields, the export carries no watermark; requesting them without a configured identity fails with the invalid code.
The anki family gains six methods that extend the spaced-repetition scheduler (FSRS): anki.reviewLog (a log of every review — rating and FSRS outcome — powering retention statistics and a faithful history), anki.forecast (a projection of the number of cards due over the coming days, including overdue cards), anki.suspendCard / anki.buryCard / anki.removeCard (remove a card from the review queue temporarily or permanently) and anki.retention (the pass rate measured over a deck’s reviews, read against the target its owner set).
Note
anki.retention replaces anki.optimizeParams, which nudged the target toward the observed rate and could write it. The retention target is a choice about the load/quality trade-off, the measured rate is its outcome, and coupling one to the other is exactly the mechanism FSRS’s authors reject. The method only measures, and never writes.
The stats family provides stats.playerTable, which returns one row of statistics per player (matches, wins/losses, counted decisions, overall / checker / cube PR, Snowie Error Rate, errors, blunders and luck) over the matches the given filter retains. As in the graphical interface, this table honours only the filter’s date range, tournaments and match lengths: the player selection and the decision type are ignored, since the table covers every player and already splits checker and cube into separate columns. The luck_known field says whether luck was measured for that player; luck_rate_mp must not be read when it is false, unknown luck not being zero luck.
The filter passed to the stats methods accepts, alongside PlayerName, a PlayerAliases field: the other spellings the same person signed under. Since a player’s name is typed by hand into every file, one person routinely appears under several spellings, and a filter that keeps only one computes on part of the matches without anything looking wrong. The field is purely additive: the decisions of any of the names are kept. Merging the names in place (MergePlayers) is the other answer, to be reserved for databases you did not receive from someone else — it rewrites everybody’s matches.
Two methods complete parity with the graphical interface: stats.tournamentBadges returns, for every tournament in the database, the figure shown on its card (the reference player’s PR), and matches.findByHash says whether a given match is already present, from the two duplicate-detection fingerprints — enough to avoid a redundant import before starting it.
The winner field of a game, received by matches.createGame and returned by matches.games, has a single encoding: 1 for player 1, -1 for player 2, 0 for an unfinished game. A client that still sends 0, 1 or -1 in gnubg’s sense (0 for player 1, 1 for player 2) records the opposite winner.
analyses.repair recomputes an analysis’s denormalised columns (cube_error among them) from its full analysis, and returns how many rows were actually corrected. Those columns are only a projection: a faulty projection therefore repairs without re-importing the source files. The operation is explicit and never triggered on its own — neither when a database is opened nor by a migration, the schema not being at fault. An unreadable analysis is left as it is rather than reset to zero. The known case: the non-doubles labelled “Double No” by gnuBG, misread before version 0.33.0, which carried the error of a double that never took place.
gammonnet.analyzeMissing triggers the gammonNet catch-up of the current tenant: write an analysis for every position that has none (ADR-0013, ADR-0015). It is a library operation — it reads and writes stored positions and analyses — never a bare evaluator: blunderdb serve operates on a library, gammonnet serve evaluates a position. The response is an NDJSON stream (started, progress, then done or error/cancelled), on the same model as the import endpoints; gammonnet.analyzeMissing.cancel (with the job_id received in the started event) cancels a running catch-up and serves either sweep indifferently — a catch-up or a re-analysis (below). It is the same operation as the automatic run after an import and the explicit gesture of the graphical interface, and as the blunderdb analyze subcommand (see Command Line Interface (CLI)) — three forms, one logic.
gammonnet.sweepStale is analyzeMissing’s counterpart for re-analysis rather than gap-filling: every position whose analysis is entirely gammonNet’s own but stale — an engine version older than the one currently running, or a depth different from ply — is re-evaluated at the requested depth. The staleness predicate is shared with the same batch in the graphical interface and in blunderdb analyze --stale (no duplicated logic across the three modes); a position carrying an XG, GNUbg or BGBlitz analysis is never touched, whatever its gammonNet content — ADR-0013’s protection remains unconditional. Same NDJSON shape as analyzeMissing, and the final event of each of the two routes carries the evaluated/refused/failed breakdown: a position gammonNet declines to evaluate (a match score beyond its table’s range, a cube decision the model refuses) counts as refused, not failed — it is never retried in vain on the next pass, unlike a position that genuinely failed.
rollout.position plays a library position (positionId) by a rollout and returns, for each candidate, the equity, its 95% interval and the JSD; rollout carries the settings (fast, standard or standard,ply=1…), store records the finished rollout as a second analysis, beside the one the position carries, which it never replaces. A bare position (an XGID) is refused: the daemon operates on a library. rollout.filter is the batch form of blunderdb analyze --rollout: the positions chosen by query (the search language) that do not yet carry a rollout with the same settings are played one after the other and recorded as they go, as an NDJSON stream (started, progress after each series of games, then done or cancelled); rollout.filter.cancel cancels it with its job_id. A tenant runs only one batch at a time, rollout or gammonNet. rollout.list reads the recorded rollouts of a position.
Correlation and business metrics
Every request gets a correlation identifier: the one the client (or a reverse proxy) sends in the X-Request-Id header, otherwise a generated one — either way returned on the same header of the response and added to the request’s closing log line (field request_id). A traceparent (W3C Trace Context), when present, is relayed as-is into that same log line — the daemon neither parses nor validates it, and embeds no tracing library: it is a bridge for correlating these logs with a tracing pipeline running upstream, nothing more.
Beyond request volume and latency, /metrics publishes gauges on the work in flight, otherwise invisible for a stuck import or gammonNet batch (one very long request, not many requests):
blunderdb_imports_inflight— imports under way, across all tenants;blunderdb_import_spool_bytes— bytes currently reserved against the import spool quota (see--rate-limit-*above for the requests-per-second counterpart);blunderdb_gammonnet_sweep_inflight— gammonNet catch-up sweeps under way, across all tenants;blunderdb_database_size_bytes— size of the main SQLite file, orpg_database_sizeunder PostgreSQL (the whole database, not per tenant, like the connection-pool gauges below); absent until a first measurement has been published.
A memory or CPU profile of the process is available by starting with --pprof-addr <host:port> (net/http/pprof): off by default, and deliberately on an address separate from --addr, since these endpoints have no notion of tenant.
Stream compression
NDJSON listings repeat the same field names on every line. The daemon compresses them when the client accepts it: send Accept-Encoding: gzip and the response comes back as Content-Encoding: gzip. Measured on a match listing: 13.5% of the original size over a thousand lines, 14.6% over a hundred.
Compression changes nothing about the stream being incremental — each record is pushed to the client as before, only compressed on the way. It applies to NDJSON, JSON and text responses only: a database export or a .dbx container is already compressed, and re-gzipping it would only make it bigger. Accept-Encoding: gzip;q=0 refuses it explicitly.
A single tenant on SQLite
The SQLite backend has no tenant column: all the data sits in the same tables, with no partition. On that backend the daemon therefore refuses any X-Tenant-ID other than 1 — accepting the others would mean serving everyone everybody’s rows behind a header claiming otherwise. A deployment that genuinely has several tenants needs the PostgreSQL backend.
Backup and restore
Four gestures, depending on what is to be recovered.
Everything, under PostgreSQL — pg_dump is the tool, and blunderDB has nothing to add:
pg_dump --format=custom --file=blunderdb.dump "postgres://…"
pg_restore --dbname="postgres://…" blunderdb.dump
Everything, under SQLite in a container — the file is opened in WAL mode (the daemon encodes journal_mode(WAL) in its connection string, for every connection in the pool): alongside blunderdb.db live a -wal and a -shm, and the most recent writes are in the -wal. Copying the .db alone from a running daemon therefore gives an incomplete file, with nothing to signal it. Two safe ways:
stop the daemon, then copy the whole volume — once stopped, the three files are consistent, and it is the volume, not the
.dbalone, that is the unit to back up;do not copy the file at all:
/v1/exports.sqlite(below) writes a complete.dbwhile the daemon is running, and it is the only gesture that requires no interruption.
/ops/maintenance.vacuum does fold the WAL back into the main file before rewriting it, but it does not freeze the database: the next write starts a new WAL. It is a compaction command, not a backup method.
One tenant alone — /v1/exports.sqlite writes a tenant’s database into an ordinary .db file, the one the desktop application opens:
curl -X POST http://127.0.0.1:8080/v1/exports.sqlite \
-H "X-Tenant-ID: 42" -o tenant-42.db
This command runs on the daemon’s own machine: it targets the local listener, bypasses the proxy, and therefore sets the tenant header itself. From outside, it is the proxy that is queried, and the tenant is that of the authenticated account — the header is not to be supplied, the proxy strips the client’s before injecting its own:
curl -u alice:… -X POST \
https://blunderdb.example.com/v1/exports.sqlite -o tenant-alice.db
Putting that file back — migrate copies it under the tenant you name:
./blunderdb migrate --from tenant-42.db --to "postgres://…" --tenant-id 42
migrate refuses to write into a tenant that already holds something, and says what (“128 positions, 3 matches”); --on-conflict skip goes ahead anyway and lets Zobrist de-duplication merge the positions.
What migrate does not copy, and announces at the end with the exact count: the Anki decks and their cards, the filter library, the search and command histories, and the session state. These are the desktop application’s usage data; the positions they refer to have indeed been moved.
The error and blunder thresholds, for their part, are copied: they are not usage data but the reading habit the counts depend on, and a tenant that counted differently from the file it came from would make the migration a silent change of meaning.
A tenant sets its own through POST /v1/librarySettings.load and /v1/librarySettings.save. Unlike metadata, which is global infrastructure exposed read-only, the settings table carries a tenant_id and lives under Row-Level Security: a tenant writing its thresholds reaches nothing but its own rows.
The desktop and the server
The desktop application opens .db files, not URLs: it never connects to any serve daemon, and there is nowhere a field to type an address. The server and the desktop exchange files, in two symmetric gestures:
from the server to the desktop —
POST /v1/exports.sqlitewrites the whole current tenant into a.dbthe desktop application opens as-is (see Backup and restore);from the desktop to the server —
blunderdb migratecopies a.dbunder the intended tenant (see Migrating a SQLite database to PostgreSQL).
There is no cross-tenant reading. The partitioning is total: nothing one tenant stores is visible to another, through any route, and no call takes a tenant as a parameter — each request knows only the one the proxy set for it. A coach who wants to see their students’ matches therefore has two paths, both explicit:
open an additional account for them in the proxy, associated with the student’s tenant: it is the proxy’s mapping table, never the daemon, that decides which tenant a session sees;
ask them for an export — the
.dbproduced byexports.sqliteor by the desktop application’s export window — and open it on their own desktop.
Deployment with Docker
The repository provides a Dockerfile.serve that builds a minimal container image of the daemon: only the serve binary is compiled (pure Go, with no graphical interface and no CGO, therefore statically linked), then placed into a distroless image.
# build
docker build -f Dockerfile.serve -t blunderdb-serve .
# run
docker run --rm -p 127.0.0.1:8080:8080 \
-e BLUNDERDB_DSN="postgres://user:pass@host:5432/blunderdb?sslmode=disable" \
blunderdb-serve
The build is run from the repository root, and the image’s default backend is postgres.
The image listens on port 8080 and is configured through environment variables (BLUNDERDB_BACKEND, BLUNDERDB_DSN, BLUNDERDB_ADDR, BLUNDERDB_RLS). It declares a HEALTHCHECK that runs blunderdb healthcheck every 30 seconds (a request on /readyz — the distroless image has neither curl nor a shell): docker ps shows the container as healthy or unhealthy, and Compose or an orchestrator can wait for the daemon to be ready before starting what depends on it.
Published image
There is no need to build the image yourself: every published blunderDB release pushes its own to the GitHub registry (GHCR), under the name ghcr.io/kevung/blunderdb-serve. Two tags are available: the version number, frozen forever on that image, and latest, which follows the most recently published version. All the documentation writes them as ghcr.io/kevung/blunderdb-serve:<version>: it is a published version number that takes the place of <version>, and it is that form, never latest, that a production deployment pins. The image is provided for linux/amd64 and linux/arm64; Docker picks the host architecture.
# pull
docker pull ghcr.io/kevung/blunderdb-serve:<version>
# postgres
docker run --rm -p 127.0.0.1:8080:8080 \
-e BLUNDERDB_DSN="postgres://user:pass@host:5432/blunderdb?sslmode=disable" \
ghcr.io/kevung/blunderdb-serve:<version>
# sqlite
docker run --rm -p 127.0.0.1:8080:8080 \
-v blunderdb-data:/data \
-e BLUNDERDB_BACKEND=sqlite -e BLUNDERDB_DSN=/data/blunderdb.db \
ghcr.io/kevung/blunderdb-serve:<version>
/data is the mount point the image prepares, with its unprivileged user’s permissions, and its XDG_DATA_HOME: the volume mounted there is not only for the database — the bearoff tables are computed there once, in /data/blunderdb, and found again on later startups. Without a volume, they are recomputed on every container start — a few seconds — and the daemon says so at startup if it cannot write them (could not prepare the bearoff tables; the exact regime will be unavailable), in which case it serves normally, with only the estimated regime on bearoff positions.
The image carries the usual OCI labels (org.opencontainers.image.source, .version, .revision, .licenses): docker inspect tells which commit and which version it comes from. It is built by continuous integration from the repo’s Dockerfile.serve, exactly as above; building locally or pulling the published image gives the same binary.
Warning
Like the daemon itself, the container performs no authentication (ADR-0005): it trusts the X-Tenant-ID header exactly as it receives it. It must be placed behind a reverse proxy responsible for authentication, which sets that header itself, and must never be exposed directly to the public Internet. The examples above publish the port on 127.0.0.1 only for this reason, and --addr likewise binds to 127.0.0.1: the proxy is on the same machine.
Deployment behind an authenticating proxy
ADR-0005 makes the reverse proxy the whole of the daemon’s security boundary: it alone authenticates the caller, it alone is allowed to set the X-Tenant-ID header, and it must strip any value sent by the client before injecting the authenticated tenant — otherwise anyone can impersonate any tenant simply by naming it. The threat model fits in one sentence: the daemon assumes a trusted internal network, and whoever reaches it directly is, as far as it is concerned, the tenant it claims to be. The repository ships a complete, ready-to-run example in the deploy/ directory. It lives in the git repository, not in the container image: so clone the repository, or download the two files reproduced below together with deploy/.env.example into the same directory.
The Compose file puts Caddy — demonstration HTTP Basic authentication — in front of blunderdb-serve and PostgreSQL, with Row-Level Security enabled. Only Caddy publishes a port: the other two services live on a Docker network declared internal: true, which has no route to either the host or the Internet, whatever ports: a later change might add to them.
# Example deployment of `blunderdb serve` behind an authenticating reverse
# proxy — the security model ADR-0005 requires and, until now, that no example
# in this repository actually showed. See deploy/README.md for the threat
# model and doc/source/mode_headless.rst for the full walkthrough.
#
# Try it from the repository root:
# POSTGRES_PASSWORD=changeme docker compose -f deploy/docker-compose.yml up -d --build
# curl -u alice:demo-password http://localhost:8080/v1/metadata.counts -d '{}'
# docker compose -f deploy/docker-compose.yml down -v
services:
# Caddy is the ENTIRE security boundary (ADR-0005): it is the only service
# with a published port, it authenticates every request, and it is the
# only thing allowed to set X-Tenant-ID — see Caddyfile. Any reverse proxy
# capable of stripping and re-setting a header works equally well; Caddy is
# used here for its one-file config and built-in Basic Auth with no extra
# modules. deploy/nginx-tenant-proxy.conf shows the equivalent nginx
# snippet for an existing nginx deployment.
caddy:
image: caddy:2-alpine
restart: unless-stopped
ports:
- "8080:80" # the ONLY port this compose project exposes to the host
volumes:
- ./Caddyfile:/etc/caddy/Caddyfile:ro
- caddy-data:/data
- caddy-config:/config
networks:
- edge # the published port lives here — "backend" is internal-only
- backend
depends_on:
blunderdb-serve:
condition: service_healthy
# No `ports:` here — on purpose (ADR-0005). The daemon performs no
# authentication of its own, so it must be reachable only from Caddy, over
# the "backend" network, and never published to the host.
blunderdb-serve:
build:
context: ..
dockerfile: Dockerfile.serve
restart: unless-stopped
environment:
BLUNDERDB_BACKEND: postgres
BLUNDERDB_DSN: "postgres://blunderdb:${POSTGRES_PASSWORD:?set POSTGRES_PASSWORD, see deploy/.env.example}@postgres:5432/blunderdb?sslmode=disable"
BLUNDERDB_ADDR: ":8080"
# Row-Level Security: defence-in-depth *inside* the trust boundary
# Caddy draws above — it does not replace the proxy (ADR-0005).
BLUNDERDB_RLS: "true"
volumes:
# The bearoff tables are computed on first start and kept under
# $XDG_DATA_HOME/blunderdb, which the image sets to /data: without a
# volume they are recomputed at every restart of the container.
- blunderdb-data:/data
networks:
- backend
depends_on:
postgres:
condition: service_healthy
postgres:
image: postgres:16-alpine
restart: unless-stopped
environment:
POSTGRES_DB: blunderdb
POSTGRES_USER: blunderdb
POSTGRES_PASSWORD: ${POSTGRES_PASSWORD:?set POSTGRES_PASSWORD, see deploy/.env.example}
volumes:
- postgres-data:/var/lib/postgresql/data
networks:
- backend
healthcheck:
test: ["CMD-SHELL", "pg_isready -U blunderdb -d blunderdb"]
interval: 5s
timeout: 3s
retries: 10
volumes:
postgres-data:
# Bearoff tables computed by blunderdb-serve on first start (see
# XDG_DATA_HOME above): a few megabytes, worth keeping across restarts.
blunderdb-data:
caddy-data:
caddy-config:
networks:
# Caddy's own network, carrying the one published port. A container on
# "backend" alone (blunderdb-serve, postgres) is never reachable through it.
edge: {}
# internal: true means this network has no route to the outside world and
# accepts no published ports — blunderdb-serve and postgres can only ever
# be reached by another container attached to it (here, only Caddy),
# never from the host or the public internet, regardless of what `ports:`
# a future edit might add to either service.
backend:
internal: true
The Caddyfile authenticates, maps the authenticated account to the tenant’s integer (map), then injects it into X-Tenant-ID after explicitly clearing any value received from the client: the header_up X-Tenant-ID "" guard precedes the injection, so that a header sent by the client can never reach the daemon, whatever later changes are made to the file.
# Demonstration reverse-proxy for `blunderdb serve` (ADR-0005).
#
# This is the WHOLE security boundary of the daemon: it authenticates the
# caller (here, HTTP Basic Auth — swap for forward_auth to a real identity
# provider, or an OIDC plugin, in production) and is the only thing allowed
# to set X-Tenant-ID. blunderdb-serve trusts that header completely and
# performs no authentication of its own.
#
# Demo credentials — CHANGE THESE before using this anywhere but a laptop:
# alice / demo-password
# bob / demo-password
# Generate a real hash with:
# docker run --rm caddy:2-alpine caddy hash-password --plaintext '<password>'
{
# This demo terminates plain HTTP on a fixed port instead of Caddy's
# automatic HTTPS, which needs a real public domain name to obtain a
# certificate for. Point a domain at this host, replace ":80" below with
# that domain, and delete these two lines to get HTTPS for free.
auto_https off
admin off
}
:80 {
basic_auth {
alice $2a$14$7yGnM3/IY8G/.mcBHMbDveecDQbnJnvHPcJQZcbWqn.H.mpttw9/.
bob $2a$14$7yGnM3/IY8G/.mcBHMbDveecDQbnJnvHPcJQZcbWqn.H.mpttw9/.
}
# Map the authenticated login (Caddy sets {http.auth.user.id} once
# basic_auth succeeds) to the tenant's positive integer — the only
# spelling of X-Tenant-ID the daemon accepts (ADR-0005, amendment
# 2026-09-03). This is the identity-to-tenant mapping ADR-0005 says is
# the proxy's job: the daemon never sees "alice", only "1".
map {http.auth.user.id} {tenant_id} {
alice 1
bob 2
default 0
}
# Never reach the daemon's operator routes or its metrics through the
# public proxy: /ops/ (vacuum, tenant purge) acts beyond the calling
# tenant, /metrics needs no X-Tenant-ID and describes the whole daemon.
@private path /ops/* /metrics
respond @private 403
reverse_proxy blunderdb-serve:8080 {
# Guard, then inject: clear whatever the client sent BEFORE setting
# the authenticated value, so a client-supplied X-Tenant-ID can never
# reach the daemon no matter how this file is edited later — the
# second line is the only one that can still be in effect once both
# have run.
header_up X-Tenant-ID ""
header_up X-Tenant-ID {tenant_id}
}
}
Two more files round out the directory: deploy/nginx-tenant-proxy.conf carries the same scheme as an nginx snippet (proxy_set_header X-Tenant-ID "" then proxy_set_header X-Tenant-ID $tenant_id, with the map $remote_user $tenant_id block), for anyone who already has an nginx in place; deploy/README.md states the threat model and what must never be done.
The Caddyfile’s HTTP Basic authentication is a demonstration, not a production recommendation: replace it with forward_auth to a real identity provider (OIDC, enterprise SSO…), which authenticates and then hands off the identity at the same spot in the file. The two passwords and the two accounts in the mapping table are to be replaced the same way.
Full scenario, from nothing to a daemon that answers:
git clone https://github.com/kevung/blunderDB.git
cd blunderDB/deploy
cp .env.example .env # POSTGRES_PASSWORD
docker compose up -d --build
# 401
curl -i http://localhost:8080/v1/metadata.counts -d '{}'
# 200
curl -u alice:demo-password http://localhost:8080/v1/metadata.counts -d '{}'
curl -u alice:demo-password -H "X-Tenant-ID: 999" \
http://localhost:8080/v1/metadata.counts -d '{}'
docker compose logs blunderdb-serve
docker compose down -v
The first request is rejected by Caddy, before it even reaches the daemon. The next two are authenticated as “alice”, whom the mapping table associates with tenant 1: they return the same body ({"positions":0,"analyses":0,"matches":0,…}) and the daemon’s log carries tenant=1 for both — the value 999 sent by the client did not survive the Caddyfile’s guard. This scenario has been replayed as written.
To pull the published image instead of building it, replace the three build: lines of the blunderdb-serve service in docker-compose.yml with a single image: line, then run docker compose up -d without --build:
blunderdb-serve:
image: ghcr.io/kevung/blunderdb-serve:<version>
restart: unless-stopped
The Compose file publishes Caddy’s port on every interface (8080:80): that is what one expects from a proxy, which is there to be reached. What must never be published is the daemon — and it is not, it has no ports: at all.
Updating a deployment
The schema is migrated automatically at startup, and this migration is one-way: a database migrated to a recent schema is no longer readable by an earlier version of blunderDB (see Annex: Database Schema). The order of the steps therefore matters.
Back up first, before anything else: it is the only way back (see Backup and restore).
Pull the tag of the intended version, never
latestin production.latestfollows the most recently published version: a deployment that pins it changes version at the whim of restarts, without anyone deciding to, and without the backup from step 1 necessarily being recent.Restart the daemon on the new image. It migrates the schema before serving a single request; if the migration fails, it stops on the error rather than serve a half-migrated database.
Check the readiness probe.
GET /readyzanswers200and{"status":"ready","version":"…"}when storage responds and its schema matches the binary’s;503and{"status":"down"}when the database is unreachable;503and{"status":"version_mismatch","version":"…","expected":"…"}when the two schemas differ — the response names the database’s and the one the binary expects.blunderdb healthcheckreturns the same verdict as an exit code.
A version_mismatch that persists after the restart means a rollback: an older binary facing an already-migrated database. There is no downward migration; it is the backup from step 1 that must be restored.
PostgreSQL backend and multi-user
For a shared deployment, blunderDB can store data in PostgreSQL rather than in a SQLite file. The backend is selected with --backend postgres and the --dsn connection string. The schema is created and migrated automatically at startup.
Data is partitioned per tenant: each request carries its tenant’s identifier (the X-Tenant-ID header, a positive decimal integer such as 1 or 42), which lets several users share the same instance without seeing each other’s data. An identifier that is not such an integer — a name like alice or default, 0, 007 — is refused with 400 invalid: the reverse-proxy maps an account to its integer, the daemon never guesses.
Row-Level Security
The --rls option additionally enables PostgreSQL’s Row-Level Security. At every startup, the daemon installs on every table carrying a tenant_id a tenant_isolation policy that lets through only the rows of the tenant named by the session parameter current_setting('app.tenant_id'), and enforces it down to the table’s owner (FORCE ROW LEVEL SECURITY). This parameter is set on the connection when it leaves the pool and reset when it returns; a connection with no tenant sees no row and inserts none. This is an optional defence in depth, disabled by default: the application code’s tenant filtering stays in place either way.
The connection role must be ordinary: neither superuser nor
BYPASSRLS. PostgreSQL lets both of those sail through every policy without a word, and isolation falls back to the application code alone. That same role must, on the other hand, own the tables, since it is the one that runs theALTER TABLEandCREATE POLICYstatements.On an already-populated database, there is nothing to migrate: applying the policies is idempotent DDL, replayed at every startup after the schema migration. No data is moved, no row rewritten; turning
--rlson or off is just a restart.The cost is measured: on reading a position, 101.8 µs without, 177.0 µs with, i.e. +73.8 % — same container, same rows, two pools differing only by this flag. It is paid on every connection borrowed from the pool (setting then resetting the parameter) and on the predicate every query crosses in addition, never on the volume of data.
Opening and closing a tenant
There is nothing to create on the server side: a tenant is not a record, it is the integer its rows carry. The database has no table of tenants and the daemon keeps no list of them — opening an account means adding an entry to the proxy’s mapping table, and the member’s first write is what makes their tenant exist.
An empty tenant answers like an empty database, with no error: metadata.counts returns zeros and the lists return nothing.
When a tenant is decommissioned, POST /ops/tenant.purge permanently deletes all its data (positions, matches, collections, history, etc.) for the current tenant (the one carried by X-Tenant-ID), as well as its session state (last search, last position, open tabs — the rows of the session_state table carrying that tenant): the operation runs in a single transaction, is idempotent (no error purging an already-empty tenant or repeating the call), and does not affect any other tenant. It erases this tenant’s rows in every table that carries one, leaving only what belongs to no one: the metadata table, including the global schema-version row, and the migration log. The purged tenant therefore becomes exactly an empty tenant again, and its integer can be reassigned. It is only available with the PostgreSQL backend — it returns an invalid error on a SQLite backend, which has no notion of tenant.
Compaction and connection pool
POST /ops/maintenance.vacuum compacts the daemon’s SQLite file — the counterpart of the “Compact database” button in the GUI and of the blunderdb vacuum command (see Command Line Interface (CLI)), with the same disk-space guard — and returns the sizes before and after (sizeBefore, sizeAfter, in bytes). It is only available with the SQLite backend; on PostgreSQL, which has no file to compact, it returns an invalid error.
The PostgreSQL connection pool is tuned via environment variable: BLUNDERDB_POSTGRES_MAX_CONNS (50 by default), BLUNDERDB_POSTGRES_MIN_CONNS (5), BLUNDERDB_POSTGRES_MAX_CONN_LIFETIME (1h), BLUNDERDB_POSTGRES_HEALTH_CHECK_PERIOD (30s), BLUNDERDB_POSTGRES_CONNECT_TIMEOUT (5s — past it, an unreachable database fails fast instead of hanging on the operating system’s TCP timeout) and BLUNDERDB_POSTGRES_MAX_CONN_IDLE_TIME (30m — a connection opened for a burst of traffic does not sit in the pool forever once the burst is over). Each value is a Go-formatted duration (5s, 30m, 1h); absent or malformed, it falls back to its default. When --metrics is on, the pool’s state is continuously exposed on /metrics: blunderdb_pg_pool_acquired (connections currently in use), _idle (available ones), _max (the configured ceiling) and _wait_count (the cumulative count of Acquire calls that had to wait for a free connection).
Migrating a SQLite database to PostgreSQL
blunderdb migrate copies a single-user SQLite database to a PostgreSQL backend, under a chosen tenant — the integer the reverse-proxy will send in X-Tenant-ID for that user — this is the path to “upload” a desktop library to a server deployment.
blunderdb migrate \
--from sqlite:///path/to/database.db \
--to "postgres://user:pass@host:5432/db?sslmode=disable" \
--tenant-id 42
# --dry-run
blunderdb migrate --from sqlite:///path/to/database.db \
--tenant-id 42 --dry-run
The migration copies the positions, their analyses and comments, the matches (games + moves), the tournaments (with their match links) and the collections (with their composition), reassigning primary and foreign keys, all within a single transaction on the destination side: the operation is atomic (a failure leaves the destination intact, just run it again). Progress and the final summary are emitted as NDJSON on standard output. If the source database is old enough to need its own in-place schema upgrade, that runs first and emits its own "schema-migration" events (phase/done/total) before the row-by-row copy begins.
Option |
Default |
Meaning |
|---|---|---|
|
– |
source SQLite database ( |
|
– |
destination PostgreSQL DSN ( |
|
– |
destination tenant, a positive decimal integer (required except in |
|
– |
counts what would be copied without writing anything |
|
|
|
Note
Application state is not (yet) migrated: Anki decks/cards, the filter library, search and command history, and session metadata. The priority is migrating the position library and the match history.
The generic call dispatcher
In addition to the historical subcommands (Command Line Interface (CLI)), blunderdb call exposes all storage operations directly, locally. It goes through the same handlers as the serve daemon: the behaviour is therefore identical to POST /v1/<family>.<method>. This is useful for scripting and integration testing.
# --list
blunderdb call --list
# read
blunderdb call metadata.counts --db database.db
blunderdb call positions.list --db database.db --json '{"limit":10}'
blunderdb call matches.get --db database.db --json '{"id":1}'
# write
blunderdb call positions.save --db database.db --json '{"position":{...}}'
blunderdb call matches.delete --db database.db --json '{"id":42}'
# a gesture of a tournament Direction, with the version a read printed
blunderdb call directions.enterResult --db database.db --if-match '…' \
--json '{"tournamentId":3,"matchId":"m7","winner":"aa"}'
Options:
Option |
Default |
Meaning |
|---|---|---|
|
– |
SQLite file (shorthand for |
|
|
|
|
|
backend connection string |
|
|
tenant, a positive decimal integer (sent as |
|
|
request body in JSON format |
|
– |
reads the request body from a file |
|
– |
lists all |
|
– |
version sent in |
call serves the transcription moves without any flag: it works on a local file, like the CLI. Every call is a fresh process, hence its own session: sessionId can be omitted, and there is no undo from one call to the next.
The JSON response (or the NDJSON stream for *.list endpoints) is written to standard output. On error, the process exits with a non-zero code and the {"error":{…}} envelope is printed to standard output so it stays parseable (for example with jq). A response carrying a Direction-Version header prints it on standard error: that is the value the next gesture passes to --if-match. call serves direction gestures without a flag, like the CLI, since it runs locally.
Tools for an AI assistant (MCP)
blunderDB embeds no language model: it offers its tools to the assistant you already use (Claude Code, Claude Desktop, a local client), through the Model Context Protocol. The assistant searches, reads and explains; blunderDB answers with its own figures.
The tools go through the same handlers as /v1 and call:
Tool |
What it returns |
|---|---|
|
counts, period covered by the matches, schema version, frequent players |
|
positions found by a search in the command-bar grammar (described in the tool), with its canonical form |
|
comments containing given words; saved searches |
|
a position, its analysis (best moves or cube), the move played and the comment |
|
the theme of the error, its cost in millipoints and the best decision |
|
neighbouring positions; reading of an XGID; legal moves; race EPC |
|
players; overall PR, checker play, cube, by phase; recurring errors |
|
matches, detail of a match, tournaments |
|
collections and their positions; study decks |
|
draws a position without its answer, then grades the answer given |
|
rollout of a library position: equity, 95% interval and JSD per candidate |
The tools only read. Four tools write — save_position, comment_position, create_collection, add_to_collection — and are offered only on request: --write locally, --mcp-write on the daemon; rollout then gains the store argument, which records the rollout beside the position’s analysis. None deletes anything.
Locally, the assistant launches blunderdb mcp on a file (see Command Line Interface (CLI)). For Claude Code:
claude mcp add blunderdb -- blunderdb mcp --db /chemin/vers/base.db
On the daemon, the same tools answer over HTTP on POST /mcp (streamable HTTP transport, sessionless). Like /v1, /mcp requires X-Tenant-ID and each tool works in that tenant; a program that embeds pkg/blunderdb/server serves it too. The daemon authenticates nobody (ADR-0005): /mcp is protected at the proxy like /v1, and --mcp-write is decided there like --direction. Every /v1 call a tool makes goes through the daemon’s whole chain again: it is logged, counted in the metrics and charged to the tenant’s rate limit, on top of the /mcp request that carries it. A tool call therefore costs several requests; none is exempted.
Like call, blunderdb mcp migrates an older database’s schema when it opens it, even without --write.