`/v1/healthz` returned a static "ok" without touching the database, so the deploy's health probe passed regardless of whether the schema the binary expects had been migrated. An api newer than its schema sailed through the probe and then failed one query at a time on whatever column was missing — the exact failure that job ordering now prevents, with nothing to catch it if that assumption ever breaks again. It now compares the applied migration version against `moments_data::expected_schema_version()`, derived from the migrations compiled into this binary via the same `sqlx::migrate!` MIGRATOR that applies them. There is no second list of expected columns to drift from the real one. schema >= expected -> 200 "ok (schema 6)" schema < expected -> 503, error-level journal line naming both versions no migrations -> 503 cannot read table -> 200 "degraded: schema unverified (...)" A newer schema than expected stays healthy: that is an api rollback under a migrated database, and this binary's queries are still satisfiable. The reverse is not. The degraded case exists because `_sqlx_migrations` is created by moments_rw and reaches moments_ro through the default privileges in asset/sql/bootstrap-moments.sql. A role provisioned before those grants would get 42501, and failing the probe over that would take a working api offline for a permissions detail. `StoreError` gains an `Inaccessible` variant so the two are told apart by SQLSTATE (42501, 42P01) rather than by sniffing message text, and the condition is loud in both the journal and the probe output. Also corrected the startup comment in moments-api: it claimed the api/worker ordering came from systemd dependencies, which cannot be true across two hosts. It comes from deploy.yml. Verified against postgres 16 with the production role split replicated (moments_rw owning the schema, moments_ro granted through bootstrap-moments.sql, plus a legacy role without those grants) and the migrations applied by the real worker binary: current schema -> 200 "ok (schema 6)"; version 6 row deleted -> 503 "schema at 5, this binary expects 6"; table emptied -> 503 "no migrations applied"; legacy role -> 200 degraded, with `curl -fsS` exiting 0 and printing the reason. First confirmed that moments_ro can in fact read `_sqlx_migrations` under the documented grants, so the normal path is the precise one. Refs #8
9.6 KiB
CLAUDE.md
This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
Project Overview
moments is a personal activity timeline and portfolio site. It ingests developer activity from multiple forges (GitHub, Gitea, Mercurial, Bugzilla), stores raw JSON payloads in PostgreSQL, and serves a React frontend showing contribution graphs, a ranked project dashboard, and a filterable activity timeline.
Architecture
Hexagonal (ports & adapters) Rust backend with a React/TypeScript frontend.
Crate Dependency Graph
moments-entities — pure types/DTOs, no DB or HTTP deps
^
moments-core — port traits (EventReader, EventWriter, EventSource, PollerStateStore)
+ presentation reshape + poller loop
^
moments-data — sole adapter: PgStore implements all core traits
+ EventSource impls (github, gitea, hg, bugzilla)
+ SQL migrations
^
moments-api — axum HTTP API binary (read-only, connects as moments_ro)
moments-worker — ingestion daemon binary (runs migrations, connects as moments_rw)
Key Design Decisions
- Raw payload storage: upstream JSON is stored verbatim in
events.payload(JSONB). Thereshape()function inmoments-core/src/presentation.rstransforms payloads intoTimelineItemat request time — no re-ingestion needed to change presentation. - Public/private gate:
events.publicboolean controls API visibility. Onlypublic = truerows are served on the detail endpoints (events,projects,activity/summary,languages/repos); the count endpoints (activity/daily,activity/hourly,sources,languages/daily) passinclude_private = true, so private work shows up as volume without leaking repo names or messages.activity/summarybridges the two: alongside the named per-repo rows it emits at most one row per period withprivate = trueand a nullsource/repo, counting that period's private-repo activity. Without it a period spent mostly in private repos read as near-idle next to a busy contribution graph. It is deliberately not split by forge — the per-period total is already derivable fromactivity/daily, a per-forge breakdown would not be (note that a?source=filter does narrow it, by design, so the numbers stay consistent with the filter the caller asked for). - Visibility reconciliation:
publicis stamped at ingest from whatever the forge reported then, and every poller is incremental — so nothing would ever revisit a repo that later flipped public ↔ private. The github and gitea sources therefore run a reconciliation pass before ingesting (reconcile_visibilityingithub_repo.rs/gitea.rs): they re-read current visibility andUPDATE events SET publicfor the whole history of each repo. It keys offevents.repo, a stored generated column (migration 0006) that derives the repo from the payload — the same expression four read queries used to each carry their own copy of. Repos already hidden are skipped (they can't leak, and staying hidden is the safe direction), and a 404 counts as private: with the user's own token, a repo still in reach answers 200 even when private. Rate limits and transient errors never flip anything. - Wire types are hand-maintained:
ui/src/api/client.tsmirrors Rust entity types manually. - Migrations: run automatically on worker startup via
sqlx::migrate!. The API binary never runs migrations.
Frontend
React 19 + Vite 6 (SWC) + TypeScript + Bootstrap 5. State/data via @tanstack/react-query. Package manager is pnpm.
Routes: / (dashboard), /activity (timeline), /project/:source/* (project detail), /blog + /blog/:slug (blog), /cv (resume).
Build & Dev Commands
Rust
cargo build --workspace # build all crates
cargo build --workspace --release # release build
cargo clippy --workspace # lint
cargo fmt --check # format check
cargo test --workspace # run tests
# Run binaries (need DATABASE_URL)
DATABASE_URL=postgres://localhost/moments cargo run -p moments-api
DATABASE_URL=postgres://localhost/moments cargo run -p moments-worker
Frontend
cd ui
pnpm install # install deps
pnpm dev # dev server on :5173 (proxies /api/* to localhost:8080)
pnpm lint # tsc --noEmit type-check
pnpm build # production build: client bundle, then prerender
The build is three steps (see ui/package.json): tsc -b → vite build (client
SPA) → pnpm run prerender (an SSR build of src/entry-server.tsx, driven by
run-prerender.mjs, that bakes one static index.html per route into ui/dist/).
The prerender fetches data at build time from VITE_API_BASE (default
https://rob.tn/api/v1) and inlines the dehydrated react-query cache as
window.__RQ_STATE__; the client hydrates it and refetches live. So a plain
curl of any route returns full content (for crawlers / AI screeners), while the
browser keeps full interactivity. Date formatting in the shared tree is pinned to
UTC + explicit field widths so SSR and client hydration match byte-for-byte.
Database
PostgreSQL with migrations in crates/moments-data/migrations/. Two roles: moments_rw (worker, full access) and moments_ro (API, SELECT-only).
API Endpoints
All under /v1/: healthz, events, sources, projects, blog, blog/{slug}, activity/daily, forge/{source}/*, og/contributions.png.
Blog posts are markdown files with YAML frontmatter (title, slug, date; optional draft/public) in the grenade/blog Gitea repo. The worker's BlogSource polls the repo (branch-tip sha as change detection) and upserts posts into events with source='blog' and occurred_at from the frontmatter date, so imported posts keep their original publish dates. The repo is the source of truth for the full set of posts: publishing, editing, renaming, and deleting are all just pushes — each poll upserts the current tree and prunes source='blog' rows that are no longer in it.
Deployment
CI-driven via Gitea Actions (.gitea/workflows/), the source of infra truth
(hosts/ports/paths live in the workflow env, not a manifest):
deploy.yml— on push tomain(or manual dispatch): lint/test gate, build the api + worker as static musl binaries (pure-rustls, so no glibc skew) and the prerendered web bundle, then deploy each component over SSH as thegitea_ciuser with scoped sudo (asset/sudoers.d/). Services run under systemd with hardened units; the api/worker reach postgres over mTLS using the host cert. The job graph is deliberately serial — binaries → worker → api → web — and each link is load-bearing:deploy-apineedsdeploy-workerbecause the worker owns migrations (it connects asmoments_rw; the api is SELECT-only and cannot run them). The two live on different hosts, so the systemd ordering the comment inmoments-api/src/main.rsdescribes cannot reach across them.deploy-workerdoesn't finish atsystemctl restart— the unit isType=simple, so that returns before migrations do. It polls the journal, scoped to the unit's currentInvocationID, for theworker startedline the worker only logs afterstore.migrate()returns./v1/healthzchecks that ordering held rather than assuming it: it compares the applied migration version against the migrations compiled into the api binary (moments_data::expected_schema_version()), so a schema behind the binary answers 503 and fails the deploy probe instead of passing and then erroring one query at a time. A newer schema is healthy (api rollback under a migrated db). A role that cannot read_sqlx_migrationsreports degraded at 200 — a missing grant says nothing about whether the api works, so it warns loudly instead of bricking the deploy.build-webneeds both deploy jobs because the prerender fetchesVITE_API_BASEat build time and bakes whatever the api is serving when it runs. In parallel, whether the crawler snapshot matched the api came down to which runner finished first; a lost race meant a stale published snapshot until the nightly refresh, with nothing detecting it.
refresh.yml— dailyschedule:(+ manual): rebuilds and redeploys only the web tier, re-baking the prerendered crawler snapshot from the current gist (CV) and activity API without bouncing the api/worker.
One-time per-host provisioning (the gitea_ci user, its authorized_keys, the
scoped sudoers drop-in) is script/infra-setup.sh, run once per host by an
operator. Gitea repo secrets: RSYNC_SSH_KEY, QUERY_GITHUB_TOKEN,
QUERY_GITEA_TOKEN (the bare GITHUB_TOKEN/GITEA_TOKEN names are reserved by
Actions, so the worker poller's tokens use the QUERY_ prefix).
Nginx reverse-proxies /api/ to the API host and serves the per-route static
files via try_files $uri $uri/ /index.html.
Both workflows render the nginx vhost through the shared script/render-site-conf.py
rather than an inline substitution per workflow. It requires every {{PLACEHOLDER}}
in asset/nginx/site.conf.tmpl to have a matching env var and refuses to emit a
file with any placeholder left unrendered — so a variable added to the template
but forgotten in one workflow's env: fails that build instead of shipping a
broken vhost to the edge. (The former per-workflow renderers drifted exactly this
way once: WEB_LISTEN reached the template and deploy.yml but not refresh.yml,
and the nightly refresh froze every reload on oolon.)