Files
tireless/script/infra-setup.sh
rob thijssen 267cb3314d
Some checks failed
deploy / deploy (push) Failing after 5m49s
fix(deploy): quote --rsync-path, and let systemd own the state directory
Run 4 failed at `ship artifacts` with `sudo: unrecognized option '--server'`.

`R="--rsync-path=sudo rsync --mkpath"` expanded unquoted as `rsync $R` splits
into three arguments — `--rsync-path=sudo`, plus a stray `rsync` that rsync
reads as a source path — so the remote end ran `sudo --server`. Use an array.
The dashboard step quoted it inline and was unaffected, which is why only half
the deploy was broken.

Every rsync destination and every sudo command in `apply system state` has now
been exercised directly against bob as gitea_ci, rather than by another six
minute round trip: seven rsync targets, sysusers, restorecon, firewalld and
daemon-reload all pass.

That surfaced the second fault. restorecon was given /var/lib/tireless, which
does not exist on a fresh host: infra-setup.sh tried to create it before
systemd-sysusers had created the account to own it, so the attempt always raced
and always lost. Declare StateDirectory=tireless on all three units instead —
systemd creates the directory, owns it as the service user and labels it — and
drop the path from restorecon, the grant, and infra-setup.

Refs #9

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013TxK1CWPkFXqdcXMJ4hVe6
2026-08-07 16:32:31 +03:00

348 lines
16 KiB
Bash
Executable File

#!/usr/bin/env bash
#
# One-time host provisioning for tireless.
#
# Run by an operator from a workstation with full sudo — NOT by CI. See
# architecture/deployment-gitea-actions.md §2. Idempotent: re-running with no
# changes is a no-op beyond file copies.
#
# Per architecture/generic.md §7 this script never suppresses errors. Where a
# command may legitimately fail (a service not yet installed), the failure is
# handled explicitly and visibly.
set -euo pipefail
APP=tireless
API_HOST="${API_HOST:-bob.hanzalova.internal}"
API_PORT="${API_PORT:-23296}"
# Ingress runs on the office proxy, not on bob (doc/plan/design.md §6.2). The
# dashboard is served from there and /v1 is reverse-proxied across the mesh, so
# the proxy needs its own (much smaller) deploy grant.
WEB_HOST="${WEB_HOST:-hanzalova.internal}"
WEB_ROOT="${WEB_ROOT:-/var/www/tireless}"
RUNNER_PUBKEY="${RUNNER_PUBKEY:-$HOME/.ssh/id_gitea_ci.pub}"
# JWK provisioner password for the internal CA, used only to mint the first
# cert (architecture/internal-tls.md §4). Never left on a host.
PROVISIONER_PW="${PROVISIONER_PW:-$HOME/.step/secrets/provisioner}"
info() { printf '\033[1;34m==>\033[0m %s\n' "$*"; }
warn() { printf '\033[1;33m warn\033[0m %s\n' "$*" >&2; }
fatal() { printf '\033[1;31mfatal\033[0m %s\n' "$*" >&2; exit 1; }
[[ -f $RUNNER_PUBKEY ]] || fatal "runner public key not found at $RUNNER_PUBKEY.
The keypair is maintained at ~/.ssh/id_gitea_ci on roosta and is shared by every
project's deploy. Copy it — do not generate a new one."
# ---------------------------------------------------------------------------
# 1. gitea_ci account, key, journal access, scoped sudoers
# ---------------------------------------------------------------------------
provision_host() {
local host="$1"
info "provisioning $host"
if ! ssh -o ConnectTimeout=5 -o BatchMode=yes "$host" true; then
warn "$host unreachable; skipping (re-run once it is back)"
return 0
fi
ssh "$host" 'sudo useradd --system --create-home --home-dir /var/lib/gitea_ci \
--shell /usr/sbin/nologin gitea_ci || echo "gitea_ci already exists"'
ssh "$host" 'sudo install -d -o gitea_ci -g gitea_ci -m 0700 /var/lib/gitea_ci/.ssh'
rsync --rsync-path 'sudo rsync' --chown gitea_ci:gitea_ci --chmod 0600 \
"$RUNNER_PUBKEY" "$host:/var/lib/gitea_ci/.ssh/authorized_keys"
ssh "$host" 'sudo usermod -aG systemd-journal gitea_ci'
# Scoped sudoers — exactly the commands the deploy runs, nothing broader.
# Named <app>_gitea_ci so other apps on this host keep their own drop-in.
ssh "$host" "sudo tee /etc/sudoers.d/${APP}_gitea_ci >/dev/null" <<SUDOERS
gitea_ci ALL=(root) NOPASSWD: /usr/bin/rsync * /usr/local/bin/tireless-api
gitea_ci ALL=(root) NOPASSWD: /usr/bin/rsync * /usr/local/bin/tireless-worker
gitea_ci ALL=(root) NOPASSWD: /usr/bin/rsync * /usr/local/bin/tireless
gitea_ci ALL=(root) NOPASSWD: /usr/bin/rsync * /etc/tireless/config.toml
gitea_ci ALL=(root) NOPASSWD: /usr/bin/rsync * /etc/sysusers.d/tireless.conf
gitea_ci ALL=(root) NOPASSWD: /usr/bin/rsync * /etc/systemd/system/tireless-api.service
gitea_ci ALL=(root) NOPASSWD: /usr/bin/rsync * /etc/systemd/system/tireless-poller.service
gitea_ci ALL=(root) NOPASSWD: /usr/bin/rsync * /etc/systemd/system/tireless-runner.service
gitea_ci ALL=(root) NOPASSWD: /usr/bin/rsync * /etc/firewalld/services/tireless-api.xml
gitea_ci ALL=(root) NOPASSWD: /usr/bin/rsync * /var/www/tireless/
gitea_ci ALL=(root) NOPASSWD: /usr/bin/systemd-sysusers
gitea_ci ALL=(root) NOPASSWD: /usr/bin/systemctl daemon-reload
gitea_ci ALL=(root) NOPASSWD: /usr/bin/systemctl restart tireless-api.service
gitea_ci ALL=(root) NOPASSWD: /usr/bin/systemctl restart tireless-poller.service
gitea_ci ALL=(root) NOPASSWD: /usr/bin/systemctl restart tireless-runner.service
gitea_ci ALL=(root) NOPASSWD: /usr/sbin/restorecon -R /usr/local/bin/tireless-api /usr/local/bin/tireless-worker /usr/local/bin/tireless /etc/tireless
gitea_ci ALL=(root) NOPASSWD: /usr/sbin/semanage port -l
gitea_ci ALL=(root) NOPASSWD: /usr/sbin/semanage port -a -t http_port_t -p tcp ${API_PORT}
gitea_ci ALL=(root) NOPASSWD: /usr/bin/firewall-cmd --reload
gitea_ci ALL=(root) NOPASSWD: /usr/bin/firewall-cmd --get-default-zone
gitea_ci ALL=(root) NOPASSWD: /usr/bin/firewall-cmd --zone=* --query-service=tireless-api
gitea_ci ALL=(root) NOPASSWD: /usr/bin/firewall-cmd --permanent --zone=* --add-service=tireless-api
gitea_ci ALL=(root) NOPASSWD: /usr/bin/firewall-cmd --zone=* --add-service=tireless-api
SUDOERS
ssh "$host" "sudo visudo -cf /etc/sudoers.d/${APP}_gitea_ci"
# ------------------------------------------------------------------------
# 2. Service account, directories, cert ACL
# ------------------------------------------------------------------------
ssh "$host" 'sudo install -d -o root -g root -m 0755 /etc/tireless'
# /var/lib/tireless is deliberately NOT created here: the service account
# does not exist until the first deploy runs systemd-sysusers, so this always
# raced. The units declare StateDirectory=tireless, which makes systemd create,
# own and label it at first start.
# The service account needs to read the host key for mTLS to Postgres (§11).
ssh "$host" 'sudo setfacl -m u:tireless:r "/etc/pki/tls/private/$(hostname -f).pem" || \
echo "deferred: tireless user does not exist yet"'
# SELinux: the API binds a non-standard port, which must be labelled before
# the first start or the bind is denied (§10).
ssh "$host" "sudo semanage port -l | grep -qE '^http_port_t.*\\b${API_PORT}\\b' \
&& echo 'port ${API_PORT} already labelled' \
|| sudo semanage port -a -t http_port_t -p tcp ${API_PORT}"
info "$host provisioned"
}
provision_host "$API_HOST"
# ---------------------------------------------------------------------------
# 3. Proxy host: web root and a deploy grant scoped to it alone
# ---------------------------------------------------------------------------
provision_web_host() {
local host="$1"
info "provisioning ingress on $host"
if ! ssh -o ConnectTimeout=5 -o BatchMode=yes "$host" true; then
warn "$host unreachable; skipping (re-run once it is back)"
return 0
fi
ssh "$host" 'sudo useradd --system --create-home --home-dir /var/lib/gitea_ci \
--shell /usr/sbin/nologin gitea_ci || echo "gitea_ci already exists"'
ssh "$host" 'sudo install -d -o gitea_ci -g gitea_ci -m 0700 /var/lib/gitea_ci/.ssh'
rsync --rsync-path 'sudo rsync' --chown gitea_ci:gitea_ci --chmod 0600 \
"$RUNNER_PUBKEY" "$host:/var/lib/gitea_ci/.ssh/authorized_keys"
ssh "$host" "sudo install -d -o root -g root -m 0755 $WEB_ROOT"
# Deliberately narrower than the API host's grant: the proxy only ever
# receives static files. It gets no systemctl, no binaries, no config.
ssh "$host" "sudo tee /etc/sudoers.d/${APP}_web_gitea_ci >/dev/null" <<SUDOERS
gitea_ci ALL=(root) NOPASSWD: /usr/bin/rsync * ${WEB_ROOT}/
gitea_ci ALL=(root) NOPASSWD: /usr/sbin/restorecon -R ${WEB_ROOT}
SUDOERS
ssh "$host" "sudo visudo -cf /etc/sudoers.d/${APP}_web_gitea_ci"
mint_service_cert "$host" "$APP"
install_vhost "$host"
info "$host ingress provisioned"
}
# ---------------------------------------------------------------------------
# Per-service internal cert (architecture/internal-tls.md §4)
#
# The step@ timer only *renews*; the first cert is minted through the JWK
# provisioner, whose password is shipped to the host only for as long as the
# mint takes and removed even when it fails.
# ---------------------------------------------------------------------------
mint_service_cert() {
local host="$1" name="$2"
local cert="/etc/nginx/tls/cert/${name}.internal.pem"
local key="/etc/nginx/tls/key/${name}.internal.pem"
local root=/etc/pki/ca-trust/source/anchors/root-internal.pem
local state
state=$(ssh "$host" "[ -f $cert ] && sudo step certificate verify $cert \
--roots $root >/dev/null 2>&1 && echo valid || echo missing")
if [[ $state == valid ]]; then
info "$name.internal cert already valid on $host"
else
[[ -f $PROVISIONER_PW ]] || fatal "provisioner password not found at $PROVISIONER_PW"
info "minting $name.internal cert on $host"
rsync -az --rsync-path='sudo rsync' --chmod=0600 \
"$PROVISIONER_PW" "$host:/tmp/${name}-provisioner"
# --san is mandatory: modern clients ignore CN, and a CN-only cert fails
# with "no alternative certificate subject name matches target hostname".
ssh "$host" "
sudo install -d -m 0755 /etc/nginx/tls/cert /etc/nginx/tls/key
rc=0
sudo step ca certificate --force \
--provisioner lair \
--provisioner-password-file /tmp/${name}-provisioner \
--ca-url https://ca.internal \
--root $root \
--san ${name}.internal \
${name}.internal $cert $key || rc=\$?
sudo rm -f /tmp/${name}-provisioner
[ \$rc -eq 0 ] || { echo 'mint failed' >&2; exit \$rc; }
sudo chown root:root $cert $key
sudo chmod 644 $cert
sudo chmod 640 $key
sudo setfacl -m u:nginx:r $key"
fi
# Harmless before the first mint — ExecCondition makes it a no-op.
ssh "$host" "sudo systemctl enable --now step@${name}.timer"
}
# ---------------------------------------------------------------------------
# nginx vhost (architecture/reverse-proxies.md §4)
# ---------------------------------------------------------------------------
install_vhost() {
local host="$1"
local conf="tireless.internal.conf"
ssh "$host" "sudo install -d -o root -g root -m 0755 $WEB_ROOT"
# A webroot not labelled httpd_sys_content_t gives nginx a 403. Deploys
# inherit the directory's type, so labelling it once here is enough.
ssh "$host" "sudo restorecon -R $WEB_ROOT"
rsync --rsync-path='sudo rsync' --chmod 0644 \
"$(dirname "$0")/../asset/nginx/tireless.hanzalova.conf" \
"$host:/etc/nginx/sites-available/$conf"
ssh "$host" "sudo ln -sfn ../sites-available/$conf /etc/nginx/sites-enabled/$conf"
# `nginx -t` parses without binding, so it cannot see a port owned across
# http{} and stream{}. It is necessary, not sufficient — hence the served-cert
# check below rather than trusting the reload (internal-tls.md §3).
ssh "$host" "sudo nginx -t" || fatal "nginx config test failed on $host"
ssh "$host" "sudo systemctl reload nginx"
}
# ---------------------------------------------------------------------------
# Split-horizon DNS, on BOTH routers
#
# A mesh host queries its own site's router, so a record added to only one
# router resolves at that site and NXDOMAINs everywhere else — including the
# operator's workstation if it is on the other site.
# ---------------------------------------------------------------------------
register_dns() {
local proxy_ip
proxy_ip=$(getent hosts "$WEB_HOST" | awk '{print $1; exit}')
[[ -n $proxy_ip ]] || fatal "cannot resolve $WEB_HOST to give the routers an address"
for site in kosherinata hanzalova; do
local cfg="$HOME/.opn-cli/${site}.yml"
if [[ ! -f $cfg ]]; then
warn "no opn-cli config for $site; register tireless.internal there by hand"
continue
fi
info "registering tireless.internal -> $proxy_ip on $site"
# Idempotent: skip the create when the override is already there, because a
# second create silently makes a *duplicate* override rather than failing,
# and two overrides for one name is a confusing thing to debug later.
#
# The listing is a padded ASCII table: `| uuid | enabled | hostname |
# domain | rr | …`, so hostname is field 4 and domain field 5 once split on
# the pipe. Trim the padding before comparing.
if opn-cli --config "$cfg" unbound host list 2>/dev/null | awk -F'|' '
NF > 5 {
h = $4; d = $5; gsub(/[[:space:]]/, "", h); gsub(/[[:space:]]/, "", d)
if (h == "tireless" && d == "internal") found = 1
}
END { exit !found }'; then
info " override already present on $site"
else
opn-cli --config "$cfg" unbound host create \
--hostname tireless --domain internal --rr A --server "$proxy_ip"
fi
# `create` only *saves* the override; Unbound keeps serving the old zone
# until the service is reconfigured. opn-cli exposes no reconfigure verb
# (it has only host/alias/domain), so call the API endpoint directly —
# otherwise the name resolves whenever Unbound next happens to reload,
# which looks like success on a slow enough check.
apply_unbound "$cfg" "$site"
done
}
# POST /api/unbound/service/reconfigure, reading credentials from the opn-cli
# config so there is one place they live.
apply_unbound() {
local cfg="$1" site="$2"
local url key secret verify
url=$(python3 -c "import yaml;print(yaml.safe_load(open('$cfg'))['url'])")
key=$(python3 -c "import yaml;print(yaml.safe_load(open('$cfg'))['api_key'])")
secret=$(python3 -c "import yaml;print(yaml.safe_load(open('$cfg'))['api_secret'])")
verify=$(python3 -c "import yaml;print(yaml.safe_load(open('$cfg')).get('ssl_verify',True))")
local curl_opts=(-fsS -u "$key:$secret" -X POST -H 'Content-Type: application/json' -d '{}')
[[ $verify == True ]] || curl_opts+=(-k)
if curl "${curl_opts[@]}" "$url/unbound/service/reconfigure" >/dev/null; then
info " unbound reconfigured on $site"
else
warn "unbound reconfigure failed on $site — apply the pending change in the
OPNsense UI, or tireless.internal will resolve only at the other site"
fi
}
provision_web_host "$WEB_HOST"
register_dns
# ---------------------------------------------------------------------------
# 4. Manual steps that cannot be automated
# ---------------------------------------------------------------------------
cat <<'MANUAL'
Remaining one-time steps (operator, on the target host):
1. Claude Code subscription login.
The OAuth flow is interactive and must be completed *as the service account*,
because Claude Code reads credentials from $HOME:
sudo -u tireless -H /usr/bin/npx -y @anthropic-ai/claude-code@2.1.220
# then: /login, and complete the browser flow
This writes /var/lib/tireless/.claude.json. The token refreshes in place,
which is why the unit declares StateDirectory=tireless (systemd creates
it, owns it as the service account, and keeps it writable).
Skip this only if you intend to run pay-as-you-go, in which case put
ANTHROPIC_API_KEY in /etc/tireless/tireless.env instead. Do not do both:
the API key silently wins, and the subscription goes unused.
2. Gitea bot account.
Create a dedicated `tireless` user on git.lair.cafe (not your own account),
generate a token scoped to issue + PR write, and put it in
/etc/tireless/tireless.env as GITEA_TOKEN (0640 root:tireless).
Then, for each repo tireless should work on:
- add `tireless` as a collaborator with write access;
- enable branch protection on the default branch, denying `tireless` push;
- confirm it can still push refs matching `tireless/*`.
The protection rule is what keeps an unattended agent from writing to main.
Verify it rather than assuming it.
3. Confirm the proxy is serving the cert that is on disk.
The vhost, its cert and the split-horizon DNS were installed above, but
`systemctl reload nginx` exits 0 whatever nginx does with the SIGHUP — and
these certs live 24 hours, so a reload that silently did not land shows up
as an expired cert tomorrow, with every file on disk looking current
(architecture/internal-tls.md §3). Verify by serial, not expiry:
served=$(echo | openssl s_client -servername tireless.internal \
-connect hanzalova.internal:443 2>/dev/null \
| openssl x509 -noout -serial)
disk=$(ssh hanzalova.internal "sudo openssl x509 -noout -serial \
-in /etc/nginx/tls/cert/tireless.internal.pem")
[ "$served" = "$disk" ] || echo "stale: $served vs $disk"
If they disagree, restart nginx rather than reloading it.
4. Postgres role and ident mapping (architecture/generic.md §5).
On magrathea AND frankie:
- create role `tireless_rw`, and a `tireless` database;
- drop /var/lib/pgsql/18/data/pg_ident.conf.d/<this-host-fqdn>.conf
containing: cert_cn <this-host-fqdn> tireless_rw
- sudo systemctl reload postgresql-18
Both servers, or a failover locks tireless out.
MANUAL
info "infra-setup complete"