taskstore: empirically derived golden path + smoke-tested pipeline #1

Merged
kgrotel merged 1 commit from taskstore-golden-path into main 2026-08-02 18:50:58 +00:00
Owner

Replaces the documentation-assembled taskstore runtime with a sequence derived by running every step in a scratch container (beads 1.1.2 + dolt 2.1.4). Full command output evidence: docs/taskstore-golden-path.md.

What differed from the documentation

  1. BEADS_DOLT_SERVER_MODE=1 means externally-managed server, not "server mode" — with it set, bd refuses to start a server even with auto-start explicitly enabled, and init fails dialing 127.0.0.1:0. Same for BEADS_DOLT_PORT/BEADS_DOLT_SERVER_PORT. This is what broke runtime attempt 3. Server mode is selected at init time (bd init --server) and persisted in .beads/metadata.json; no BEADS_DOLT_* env may ever be set.
  2. bd's daemon management needs procps (pgrep/ps), documented nowhere — python:3.12-slim ships neither, and without them bd deletes its own PID/port state files while the server keeps running (status shows "not running, port 0", clients dial port 0, a second start collides with the dolt lock).
  3. A failed bd init leaves a partial .beads/ behind — so [ -d .beads ] is not an init guard. The only valid marker is metadata.json with "dolt_mode": "server".
  4. The dolt port is derived fresh per boot (observed 36139 → 45877 → 37089) and resolved via .beads/dolt-server.port — pinning 3307 anywhere both mismatches and triggers divergence 1.
  5. --prefix caiman must be passed at first init — without it the prefix silently defaults to the directory name (workspace); the documented "demands a prefix" error is not reliable.
  6. bd dolt status exits 0 even when the server is down — the health gate is a real SQL roundtrip (bd list --json), which fails hard exactly when it should because sessions run with dolt.auto-start: false.

Auto-start verdict (both patterns tested): hybrid. Init needs auto-start ON (it bootstraps the server it initializes against). Immediately after, dolt.auto-start: false is written to .beads/config.yaml: from then on clients (kmcp-spawned sessions included) connect or fail loudly, and the entrypoint's explicit idempotent bd dolt start is the only sanctioned daemon launch. Verified: repeat bd dolt start is a no-op; a raw second dolt sql-server is refused by the directory lock without corruption.

Migration note — existing PVC

The current PVC holds exactly one failed scratch init (partial .beads/ without server-mode metadata — reproduced and tested against). Wiping it is acceptable and is what the entrypoint does automatically: the guard detects the partial state, logs wiping for reinit, and reinitializes server-mode with prefix caiman. No manual PVC action needed; no data exists to migrate.

Rollout

  1. Merge → factory builds :x.y-rc, runs the in-image smoke test (init → status/where → create/ready/show → restart persistence → second-writer → MCP handshake), and only then promotes to the version tag + latest and prints the digest.
  2. Deploy repo PR (companion): connect-only env for the MCPServer manifest.
  3. Roll by deleting the pod, never by scaling: kubectl -n caiman delete pod -l app.kubernetes.io/name=caiman-taskstore. A second replica would block on (and must never share) the dolt lock.
  4. Pin the printed digest in base/caiman/taskstore-mcpserver.yaml.

Validation done locally (rootless podman)

  • fresh volume: clean init, MCP handshake serves tool list on stdout, setup log on stderr
  • replica of the broken PVC: wiped + reinitialized + healthy
  • restart on initialized volume: no wipe, no re-init, issue persisted
  • full smoke-test: all six steps green

🤖 Generated with Claude Code

Replaces the documentation-assembled taskstore runtime with a sequence derived by **running every step** in a scratch container (beads 1.1.2 + dolt 2.1.4). Full command output evidence: `docs/taskstore-golden-path.md`. ## What differed from the documentation 1. **`BEADS_DOLT_SERVER_MODE=1` means *externally-managed* server, not "server mode"** — with it set, bd refuses to start a server even with auto-start explicitly enabled, and init fails dialing `127.0.0.1:0`. Same for `BEADS_DOLT_PORT`/`BEADS_DOLT_SERVER_PORT`. This is what broke runtime attempt 3. Server mode is selected at init time (`bd init --server`) and persisted in `.beads/metadata.json`; **no `BEADS_DOLT_*` env may ever be set**. 2. **bd's daemon management needs `procps` (pgrep/ps), documented nowhere** — python:3.12-slim ships neither, and without them bd deletes its own PID/port state files while the server keeps running (status shows "not running, port 0", clients dial port 0, a second start collides with the dolt lock). 3. **A failed `bd init` leaves a partial `.beads/` behind** — so `[ -d .beads ]` is not an init guard. The only valid marker is `metadata.json` with `"dolt_mode": "server"`. 4. **The dolt port is derived fresh per boot** (observed 36139 → 45877 → 37089) and resolved via `.beads/dolt-server.port` — pinning 3307 anywhere both mismatches and triggers divergence 1. 5. **`--prefix caiman` must be passed at first init** — without it the prefix silently defaults to the directory name (`workspace`); the documented "demands a prefix" error is not reliable. 6. **`bd dolt status` exits 0 even when the server is down** — the health gate is a real SQL roundtrip (`bd list --json`), which fails hard exactly when it should because sessions run with `dolt.auto-start: false`. **Auto-start verdict (both patterns tested):** hybrid. Init needs auto-start ON (it bootstraps the server it initializes against). Immediately after, `dolt.auto-start: false` is written to `.beads/config.yaml`: from then on clients (kmcp-spawned sessions included) connect or fail loudly, and the entrypoint's explicit idempotent `bd dolt start` is the only sanctioned daemon launch. Verified: repeat `bd dolt start` is a no-op; a raw second `dolt sql-server` is refused by the directory lock without corruption. ## Migration note — existing PVC The current PVC holds exactly one **failed scratch init** (partial `.beads/` without server-mode metadata — reproduced and tested against). **Wiping it is acceptable and is what the entrypoint does automatically**: the guard detects the partial state, logs `wiping for reinit`, and reinitializes server-mode with prefix `caiman`. No manual PVC action needed; no data exists to migrate. ## Rollout 1. Merge → factory builds `:x.y-rc`, runs the in-image smoke test (init → status/where → create/ready/show → restart persistence → second-writer → MCP handshake), and only then promotes to the version tag + `latest` and prints the digest. 2. Deploy repo PR (companion): connect-only env for the MCPServer manifest. 3. **Roll by deleting the pod, never by scaling**: `kubectl -n caiman delete pod -l app.kubernetes.io/name=caiman-taskstore`. A second replica would block on (and must never share) the dolt lock. 4. Pin the printed digest in `base/caiman/taskstore-mcpserver.yaml`. ## Validation done locally (rootless podman) - fresh volume: clean init, MCP handshake serves tool list on stdout, setup log on stderr - replica of the broken PVC: wiped + reinitialized + healthy - restart on initialized volume: no wipe, no re-init, issue persisted - full `smoke-test`: all six steps green 🤖 Generated with [Claude Code](https://claude.com/claude-code)
Derived by running every lifecycle step in a scratch container (beads 1.1.2,
dolt 2.1.4) instead of assembling from documentation. Full evidence in
docs/taskstore-golden-path.md.

Image:
- procps added: bd's daemon tracking shells out to pgrep/ps; without them it
  deletes its own PID/port state while the server runs (root cause of the
  untracked-daemon failures)
- all BEADS_DOLT_* env REMOVED from the image: on 1.1.2 those vars select
  externally-managed server mode and break init and port resolution; server
  mode is persisted at init time, connect-only comes from dolt.auto-start
- env contract: HOME, BEADS_DIR, BEADS_WORKING_DIR only

Entrypoint (split into setup.sh + thin entrypoint):
- init only when .beads/metadata.json shows dolt_mode=server; anything else
  (partial init — the current PVC state — or embedded) is wiped and reinit'd
- bd init --server --prefix caiman with auto-start on (init bootstraps its
  own server), then dolt.auto-start: false locks sessions to connect-only
- explicit bd dolt start is the single sanctioned daemon launch; health gate
  is a real SQL roundtrip (bd list --json), not bd dolt status (exits 0 even
  when down); any failure exits non-zero

CI:
- pipeline is now build (kaniko, :version-rc) -> smoke (runs the baked-in
  /usr/local/bin/smoke-test inside the rc image) -> promote (crane, rc ->
  version + latest, prints digest); a failed lifecycle never updates tags

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
kgrotel deleted branch taskstore-golden-path 2026-08-02 21:10:03 +00:00
Sign in to join this conversation.
No reviewers
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
ccio/imagefactory!1
No description provided.