`build-web` had no `needs:`, so it ran in parallel with the binary build and both deploy jobs. The prerender fetches VITE_API_BASE at build time, which means the job baked whatever the api happened to be serving at the moment it ran — old binary or new, depending on which runner finished first. Both outcomes showed up on consecutive runs. Run 68 won by ~25 seconds and the repo-visibility fix reached the api and the crawler snapshot together. Run 69 lost: the api served the new activity/summary private aggregate while the baked /activity/ snapshot had none, so the 15 August card read "47 changes in 3 repositories" where the api said 54 with 7 private. Browsers hydrate and refetch, so a visitor saw the right page — but curl, which is what crawlers and AI screeners get, saw the stale one until refresh.yml was dispatched by hand. Nothing in the pipeline noticed. deploy-worker is in the needs list alongside deploy-api because the worker owns migrations: a schema change isn't live until it has restarted, so an api that depends on one isn't answering correctly before then either. The graph is now serial — binaries -> api/worker -> web — which costs the overlap between the web build and the binary build. Worth it: the alternative is a published snapshot whose correctness depends on runner scheduling, failing silently and for a whole day. Job graph verified acyclic with every `needs:` resolving to a real job. Closes #8
8.5 KiB
CLAUDE.md
This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
Project Overview
moments is a personal activity timeline and portfolio site. It ingests developer activity from multiple forges (GitHub, Gitea, Mercurial, Bugzilla), stores raw JSON payloads in PostgreSQL, and serves a React frontend showing contribution graphs, a ranked project dashboard, and a filterable activity timeline.
Architecture
Hexagonal (ports & adapters) Rust backend with a React/TypeScript frontend.
Crate Dependency Graph
moments-entities — pure types/DTOs, no DB or HTTP deps
^
moments-core — port traits (EventReader, EventWriter, EventSource, PollerStateStore)
+ presentation reshape + poller loop
^
moments-data — sole adapter: PgStore implements all core traits
+ EventSource impls (github, gitea, hg, bugzilla)
+ SQL migrations
^
moments-api — axum HTTP API binary (read-only, connects as moments_ro)
moments-worker — ingestion daemon binary (runs migrations, connects as moments_rw)
Key Design Decisions
- Raw payload storage: upstream JSON is stored verbatim in
events.payload(JSONB). Thereshape()function inmoments-core/src/presentation.rstransforms payloads intoTimelineItemat request time — no re-ingestion needed to change presentation. - Public/private gate:
events.publicboolean controls API visibility. Onlypublic = truerows are served on the detail endpoints (events,projects,activity/summary,languages/repos); the count endpoints (activity/daily,activity/hourly,sources,languages/daily) passinclude_private = true, so private work shows up as volume without leaking repo names or messages.activity/summarybridges the two: alongside the named per-repo rows it emits at most one row per period withprivate = trueand a nullsource/repo, counting that period's private-repo activity. Without it a period spent mostly in private repos read as near-idle next to a busy contribution graph. It is deliberately not split by forge — the per-period total is already derivable fromactivity/daily, a per-forge breakdown would not be (note that a?source=filter does narrow it, by design, so the numbers stay consistent with the filter the caller asked for). - Visibility reconciliation:
publicis stamped at ingest from whatever the forge reported then, and every poller is incremental — so nothing would ever revisit a repo that later flipped public ↔ private. The github and gitea sources therefore run a reconciliation pass before ingesting (reconcile_visibilityingithub_repo.rs/gitea.rs): they re-read current visibility andUPDATE events SET publicfor the whole history of each repo. It keys offevents.repo, a stored generated column (migration 0006) that derives the repo from the payload — the same expression four read queries used to each carry their own copy of. Repos already hidden are skipped (they can't leak, and staying hidden is the safe direction), and a 404 counts as private: with the user's own token, a repo still in reach answers 200 even when private. Rate limits and transient errors never flip anything. - Wire types are hand-maintained:
ui/src/api/client.tsmirrors Rust entity types manually. - Migrations: run automatically on worker startup via
sqlx::migrate!. The API binary never runs migrations.
Frontend
React 19 + Vite 6 (SWC) + TypeScript + Bootstrap 5. State/data via @tanstack/react-query. Package manager is pnpm.
Routes: / (dashboard), /activity (timeline), /project/:source/* (project detail), /blog + /blog/:slug (blog), /cv (resume).
Build & Dev Commands
Rust
cargo build --workspace # build all crates
cargo build --workspace --release # release build
cargo clippy --workspace # lint
cargo fmt --check # format check
cargo test --workspace # run tests
# Run binaries (need DATABASE_URL)
DATABASE_URL=postgres://localhost/moments cargo run -p moments-api
DATABASE_URL=postgres://localhost/moments cargo run -p moments-worker
Frontend
cd ui
pnpm install # install deps
pnpm dev # dev server on :5173 (proxies /api/* to localhost:8080)
pnpm lint # tsc --noEmit type-check
pnpm build # production build: client bundle, then prerender
The build is three steps (see ui/package.json): tsc -b → vite build (client
SPA) → pnpm run prerender (an SSR build of src/entry-server.tsx, driven by
run-prerender.mjs, that bakes one static index.html per route into ui/dist/).
The prerender fetches data at build time from VITE_API_BASE (default
https://rob.tn/api/v1) and inlines the dehydrated react-query cache as
window.__RQ_STATE__; the client hydrates it and refetches live. So a plain
curl of any route returns full content (for crawlers / AI screeners), while the
browser keeps full interactivity. Date formatting in the shared tree is pinned to
UTC + explicit field widths so SSR and client hydration match byte-for-byte.
Database
PostgreSQL with migrations in crates/moments-data/migrations/. Two roles: moments_rw (worker, full access) and moments_ro (API, SELECT-only).
API Endpoints
All under /v1/: healthz, events, sources, projects, blog, blog/{slug}, activity/daily, forge/{source}/*, og/contributions.png.
Blog posts are markdown files with YAML frontmatter (title, slug, date; optional draft/public) in the grenade/blog Gitea repo. The worker's BlogSource polls the repo (branch-tip sha as change detection) and upserts posts into events with source='blog' and occurred_at from the frontmatter date, so imported posts keep their original publish dates. The repo is the source of truth for the full set of posts: publishing, editing, renaming, and deleting are all just pushes — each poll upserts the current tree and prunes source='blog' rows that are no longer in it.
Deployment
CI-driven via Gitea Actions (.gitea/workflows/), the source of infra truth
(hosts/ports/paths live in the workflow env, not a manifest):
deploy.yml— on push tomain(or manual dispatch): lint/test gate, build the api + worker as static musl binaries (pure-rustls, so no glibc skew) and the prerendered web bundle, then deploy each component over SSH as thegitea_ciuser with scoped sudo (asset/sudoers.d/). Services run under systemd with hardened units; the api/worker reach postgres over mTLS using the host cert. The job graph is deliberately serial — binaries → api/worker → web — because the prerender fetchesVITE_API_BASEat build time and so bakes whatever the api is serving whenbuild-webruns.build-webthereforeneeds:both deploy jobs (the worker owns migrations, so a schema change isn't live until it restarts). When it ran in parallel instead, whether the crawler snapshot matched the api came down to which runner finished first, and a lost race meant a stale snapshot until the nightly refresh.refresh.yml— dailyschedule:(+ manual): rebuilds and redeploys only the web tier, re-baking the prerendered crawler snapshot from the current gist (CV) and activity API without bouncing the api/worker.
One-time per-host provisioning (the gitea_ci user, its authorized_keys, the
scoped sudoers drop-in) is script/infra-setup.sh, run once per host by an
operator. Gitea repo secrets: RSYNC_SSH_KEY, QUERY_GITHUB_TOKEN,
QUERY_GITEA_TOKEN (the bare GITHUB_TOKEN/GITEA_TOKEN names are reserved by
Actions, so the worker poller's tokens use the QUERY_ prefix).
Nginx reverse-proxies /api/ to the API host and serves the per-route static
files via try_files $uri $uri/ /index.html.
Both workflows render the nginx vhost through the shared script/render-site-conf.py
rather than an inline substitution per workflow. It requires every {{PLACEHOLDER}}
in asset/nginx/site.conf.tmpl to have a matching env var and refuses to emit a
file with any placeholder left unrendered — so a variable added to the template
but forgotten in one workflow's env: fails that build instead of shipping a
broken vhost to the edge. (The former per-workflow renderers drifted exactly this
way once: WEB_LISTEN reached the template and deploy.yml but not refresh.yml,
and the nightly refresh froze every reload on oolon.)