Files
moments/CLAUDE.md
rob thijssen 7e186f76ae
Some checks failed
deploy / Build prerendered web (push) Blocked by required conditions
deploy / Build api + worker (static musl) (push) Successful in 5m33s
deploy / Deploy moments-worker to frootmig (push) Successful in 18s
deploy / Deploy moments-api to nikola (push) Successful in 22s
deploy / Deploy web to oolon (push) Has been cancelled
fix(ci): restart the api only after the worker has finished migrating
The worker owns migrations — it connects as moments_rw, while the api is
SELECT-only and would fail with `permission denied for schema public` if
it tried. Nothing ordered the two deploy jobs, though, and they run
against different hosts (nikola, frootmig), so the systemd ordering the
comment in moments-api/src/main.rs appeals to cannot reach between them.

On the run that shipped the `events.repo` migration, the worker happened
to restart 5 seconds before the api. Reversed, the api would have come up
querying a column that did not exist yet and errored on every request
touching it — and `/v1/healthz` returns a static "ok" without touching
the database, so the deploy's own health probe would have passed it as
fine.

deploy-api now needs deploy-worker. That ordering is only worth anything
if the worker job stays open until migrations are actually done, and it
didn't: the unit is Type=simple, so `systemctl restart` returns once the
process is exec'd and `is-active` is true immediately, neither of which
says anything about migrations. So the job now waits for the worker to
log "worker started", which it does only after awaiting store.migrate().
The search is scoped to the unit's current InvocationID so a previous
start's line cannot satisfy the wait, and it gives up loudly after 120s
or as soon as the process exits.

The wait deliberately avoids `journalctl | grep -q`: a consumer that
exits on first match closes the pipe, journalctl dies of SIGPIPE, and
under `pipefail` the match reads as a failure — which is exactly what the
first draft did, and it would have failed every deploy on a worker that
started perfectly. journalctl --grep does the filtering instead and the
result is tested for emptiness, with no pipe in the pipeline.

Verified by extracting the gate out of the workflow and running it
against fake systemctl/journalctl: started -> exit 0 on the first check;
process gone -> exit 1 immediately with "worker exited before reporting
startup"; no InvocationID -> exit 1 immediately; never logs but stays
active -> loops and fails on timeout. Job graph re-checked acyclic with
every `needs:` resolving to a real job (binaries -> worker -> api -> web).

Closes #8
2026-08-17 13:04:19 +03:00

9.0 KiB

CLAUDE.md

This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.

Project Overview

moments is a personal activity timeline and portfolio site. It ingests developer activity from multiple forges (GitHub, Gitea, Mercurial, Bugzilla), stores raw JSON payloads in PostgreSQL, and serves a React frontend showing contribution graphs, a ranked project dashboard, and a filterable activity timeline.

Architecture

Hexagonal (ports & adapters) Rust backend with a React/TypeScript frontend.

Crate Dependency Graph

moments-entities   — pure types/DTOs, no DB or HTTP deps
       ^
moments-core       — port traits (EventReader, EventWriter, EventSource, PollerStateStore)
                     + presentation reshape + poller loop
       ^
moments-data       — sole adapter: PgStore implements all core traits
                     + EventSource impls (github, gitea, hg, bugzilla)
                     + SQL migrations
       ^
moments-api        — axum HTTP API binary (read-only, connects as moments_ro)
moments-worker     — ingestion daemon binary (runs migrations, connects as moments_rw)

Key Design Decisions

  • Raw payload storage: upstream JSON is stored verbatim in events.payload (JSONB). The reshape() function in moments-core/src/presentation.rs transforms payloads into TimelineItem at request time — no re-ingestion needed to change presentation.
  • Public/private gate: events.public boolean controls API visibility. Only public = true rows are served on the detail endpoints (events, projects, activity/summary, languages/repos); the count endpoints (activity/daily, activity/hourly, sources, languages/daily) pass include_private = true, so private work shows up as volume without leaking repo names or messages. activity/summary bridges the two: alongside the named per-repo rows it emits at most one row per period with private = true and a null source/repo, counting that period's private-repo activity. Without it a period spent mostly in private repos read as near-idle next to a busy contribution graph. It is deliberately not split by forge — the per-period total is already derivable from activity/daily, a per-forge breakdown would not be (note that a ?source= filter does narrow it, by design, so the numbers stay consistent with the filter the caller asked for).
  • Visibility reconciliation: public is stamped at ingest from whatever the forge reported then, and every poller is incremental — so nothing would ever revisit a repo that later flipped public ↔ private. The github and gitea sources therefore run a reconciliation pass before ingesting (reconcile_visibility in github_repo.rs / gitea.rs): they re-read current visibility and UPDATE events SET public for the whole history of each repo. It keys off events.repo, a stored generated column (migration 0006) that derives the repo from the payload — the same expression four read queries used to each carry their own copy of. Repos already hidden are skipped (they can't leak, and staying hidden is the safe direction), and a 404 counts as private: with the user's own token, a repo still in reach answers 200 even when private. Rate limits and transient errors never flip anything.
  • Wire types are hand-maintained: ui/src/api/client.ts mirrors Rust entity types manually.
  • Migrations: run automatically on worker startup via sqlx::migrate!. The API binary never runs migrations.

Frontend

React 19 + Vite 6 (SWC) + TypeScript + Bootstrap 5. State/data via @tanstack/react-query. Package manager is pnpm.

Routes: / (dashboard), /activity (timeline), /project/:source/* (project detail), /blog + /blog/:slug (blog), /cv (resume).

Build & Dev Commands

Rust

cargo build --workspace              # build all crates
cargo build --workspace --release    # release build
cargo clippy --workspace             # lint
cargo fmt --check                    # format check
cargo test --workspace               # run tests

# Run binaries (need DATABASE_URL)
DATABASE_URL=postgres://localhost/moments cargo run -p moments-api
DATABASE_URL=postgres://localhost/moments cargo run -p moments-worker

Frontend

cd ui
pnpm install                         # install deps
pnpm dev                             # dev server on :5173 (proxies /api/* to localhost:8080)
pnpm lint                            # tsc --noEmit type-check
pnpm build                           # production build: client bundle, then prerender

The build is three steps (see ui/package.json): tsc -bvite build (client SPA) → pnpm run prerender (an SSR build of src/entry-server.tsx, driven by run-prerender.mjs, that bakes one static index.html per route into ui/dist/). The prerender fetches data at build time from VITE_API_BASE (default https://rob.tn/api/v1) and inlines the dehydrated react-query cache as window.__RQ_STATE__; the client hydrates it and refetches live. So a plain curl of any route returns full content (for crawlers / AI screeners), while the browser keeps full interactivity. Date formatting in the shared tree is pinned to UTC + explicit field widths so SSR and client hydration match byte-for-byte.

Database

PostgreSQL with migrations in crates/moments-data/migrations/. Two roles: moments_rw (worker, full access) and moments_ro (API, SELECT-only).

API Endpoints

All under /v1/: healthz, events, sources, projects, blog, blog/{slug}, activity/daily, forge/{source}/*, og/contributions.png.

Blog posts are markdown files with YAML frontmatter (title, slug, date; optional draft/public) in the grenade/blog Gitea repo. The worker's BlogSource polls the repo (branch-tip sha as change detection) and upserts posts into events with source='blog' and occurred_at from the frontmatter date, so imported posts keep their original publish dates. The repo is the source of truth for the full set of posts: publishing, editing, renaming, and deleting are all just pushes — each poll upserts the current tree and prunes source='blog' rows that are no longer in it.

Deployment

CI-driven via Gitea Actions (.gitea/workflows/), the source of infra truth (hosts/ports/paths live in the workflow env, not a manifest):

  • deploy.yml — on push to main (or manual dispatch): lint/test gate, build the api + worker as static musl binaries (pure-rustls, so no glibc skew) and the prerendered web bundle, then deploy each component over SSH as the gitea_ci user with scoped sudo (asset/sudoers.d/). Services run under systemd with hardened units; the api/worker reach postgres over mTLS using the host cert. The job graph is deliberately serial — binaries → worker → api → web — and each link is load-bearing:
    • deploy-api needs deploy-worker because the worker owns migrations (it connects as moments_rw; the api is SELECT-only and cannot run them). The two live on different hosts, so the systemd ordering the comment in moments-api/src/main.rs describes cannot reach across them.
    • deploy-worker doesn't finish at systemctl restart — the unit is Type=simple, so that returns before migrations do. It polls the journal, scoped to the unit's current InvocationID, for the worker started line the worker only logs after store.migrate() returns.
    • build-web needs both deploy jobs because the prerender fetches VITE_API_BASE at build time and bakes whatever the api is serving when it runs. In parallel, whether the crawler snapshot matched the api came down to which runner finished first; a lost race meant a stale published snapshot until the nightly refresh, with nothing detecting it.
  • refresh.yml — daily schedule: (+ manual): rebuilds and redeploys only the web tier, re-baking the prerendered crawler snapshot from the current gist (CV) and activity API without bouncing the api/worker.

One-time per-host provisioning (the gitea_ci user, its authorized_keys, the scoped sudoers drop-in) is script/infra-setup.sh, run once per host by an operator. Gitea repo secrets: RSYNC_SSH_KEY, QUERY_GITHUB_TOKEN, QUERY_GITEA_TOKEN (the bare GITHUB_TOKEN/GITEA_TOKEN names are reserved by Actions, so the worker poller's tokens use the QUERY_ prefix). Nginx reverse-proxies /api/ to the API host and serves the per-route static files via try_files $uri $uri/ /index.html.

Both workflows render the nginx vhost through the shared script/render-site-conf.py rather than an inline substitution per workflow. It requires every {{PLACEHOLDER}} in asset/nginx/site.conf.tmpl to have a matching env var and refuses to emit a file with any placeholder left unrendered — so a variable added to the template but forgotten in one workflow's env: fails that build instead of shipping a broken vhost to the edge. (The former per-workflow renderers drifted exactly this way once: WEB_LISTEN reached the template and deploy.yml but not refresh.yml, and the nightly refresh froze every reload on oolon.)