21 Commits

Author SHA1 Message Date
rob thijssen
98f193d16b feat(data): implement JobStore and the Gitea read client
All checks were successful
deploy / deploy (push) Successful in 6m1s
Closes #3 and #4.

Both specs had gaps that only appeared on contact, which is worth recording
because it is evidence for the plan-quality question in #10:

- `enqueue(&[DiscoveredIssue])` cannot know a JobKind. It is derivable from
  labels, so the store now carries the LabelProtocol and core gained
  `job_kind_for`. The precedence when an operator applies several mode labels
  had to be decided: plan beats implement, because planning produces the
  implementation children and so loses nothing, while the reverse silently
  discards the decomposition that was also asked for.
- `claim_next(worker, allowed_lanes)` had no lane to filter on. Migration 0002
  adds one, recorded at enqueue by `routing::lane_for` — the same function
  `route` now delegates to, so the two cannot drift. Deriving it at claim time
  instead would mean a forge request per claim, since labels are not stored.
- `list_opted_in_issues` returned a bare Vec with nowhere to put the ETag that
  #4's own step 3 requires, so a caller could not make the next poll
  conditional. It returns an `IssuePage` now, which also carries `not_modified`
  — a 304 is not an empty repo, and a poller that conflated them would treat
  every quiet poll as every issue having disappeared.

The claim is one statement: a CTE takes the row lock with SKIP LOCKED and the
update writes the claim, so select and update share a transaction without
managing one by hand. Returning a job to pending clears the claim, because
`pending_holds_no_claim` refuses the half-done version — the database is what
makes lease expiry safe rather than the code remembering to.

Enum values round-trip through serde rather than a hand-written match, so the
schema's check constraints and the Rust types are provably one vocabulary.

Also replaced a test from #2 that asserted exactly one migration exists. It
failed on the first legitimate migration, which teaches people to edit the
assertion rather than think. It now asserts what it was reaching for: versions
unique and ascending.

Verified against Postgres 18 — 15 database tests covering enqueue idempotency
across repeated polls, lane filtering including the agent:oc override, claim
metadata, lease expiry and reclaim, terminal jobs refusing transition, renewal
requiring you still hold the claim, and re-routing a queued job when its labels
change. Plus 12 mock-forge tests: If-None-Match sent and 304 distinguished from
empty, ETag surfaced, 429 retried but bounded, 503 retried then succeeding, 404
not retried, pull requests filtered out, and every unimplemented write failing
without touching the forge.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013TxK1CWPkFXqdcXMJ4hVe6
2026-08-07 20:20:09 +03:00
rob thijssen
6545d1980b feat(data): add the initial schema and migrations
All checks were successful
deploy / deploy (push) Successful in 5m35s
Closes #2. Three tables mirroring tireless-entities, plus the sqlx
compile-time-checking decision the rest of stage 1 inherits.

RUNTIME QUERIES, NOT `query!`. No .sqlx metadata, no DATABASE_URL to build.
lairball, the other house project on this cluster, does the same. The deciding
argument is specific to tireless: compile-time checking makes a database a build
dependency, and this crate is meant to be modified by a 27B model working
unattended. A build that fails without a database it cannot provision is one
that model cannot fix, and its documented failure mode is to improvise. The cost
is named in store.rs — a malformed query is caught by a test, not by cargo.

Enums are text with a check constraint, not Postgres enum types: adding a
JobKind variant should be a migration, not an ALTER TYPE holding a lock. Unit
tests assert every serde variant appears in the schema, so adding a variant
without a migration fails the build rather than the first job of that kind.

That surfaced a spelling that would have been permanent: `rename_all =
"snake_case"` turns Forge::GitHub into `git_hub`, across the database, the JSON
API and the generated TypeScript. Renamed to `github` now, while nothing is
persisted and GitHub support is still disabled.

The live-issue index is PARTIAL, and both directions of getting it wrong are
silent. A plain unique constraint on (forge, owner, repo, number) would forbid
re-running a terminal job, and would forbid the discovery lane outright, since
discovery recurs against one tracking issue on a cooldown. Partial on the
non-terminal states gives at most one live job per issue and unlimited history.

The claim index orders by created_at alone and leaves kind as a filter, because
the claim takes LIMIT 1 and can stop at the first match. Leading with kind sorts
every pending row on every claim: measured at 720 buffers versus 4, and the gap
grows with the backlog rather than staying fixed. INCLUDE (kind) was measured
too and dropped — FOR UPDATE visits the heap regardless.

Verified against Postgres 18, the same major as the house cluster: migrations
apply to an empty database and are a no-op on the second run; all eight
constraints reject what they should and admit what they should; and two
concurrent claimers of one pending job produce exactly one winner, with the
loser skipping rather than blocking.

Those live tests are #[ignore]d, not skipped on a missing variable, so a green
`cargo test` never implies the schema was exercised. CLAUDE.md says how to run
them and that an applied migration must never be edited.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013TxK1CWPkFXqdcXMJ4hVe6
2026-08-07 18:47:16 +03:00
rob thijssen
9deeaaa9c2 feat(infra): provision the poller identity
All checks were successful
deploy / deploy (push) Successful in 5m35s
Stage 2 writes the first label, and nothing about the label protocol can be
dogfooded until an identity exists that is allowed to. Provisioned now rather
than deferred.

  account   tireless-poller (tireless-poller@lair.cafe), no ssh key — it never
            clones anything
  team      lair/tireless-poller: Issues=write, PullRequests=read, Code=none,
            everything else none, scoped to named repos
  token     write:issue, read:repository, read:user

This is the one place an org team earns its keep. Collaborator permissions in
Gitea are repo-wide, so granting issue write that way would hand this account
code push too — the unit-level team grants labelling without it.

The units now load different environment files, and that is not tidiness. The
runner spawns coding agents as subprocesses, and subprocesses inherit the
environment, so a label-capable token in the runner's environment is a token
every agent run can read. poller.env is loaded only by tireless-poller.service,
which spawns nothing.

Be precise about what that buys: the separation is between processes, not users.
Both units run as `tireless`, so the runner's uid can still read poller.env even
though its process never loads it. An agent would have to go looking rather than
find it handed over in `env` — a real difference, not a boundary. Closing it
means separate service accounts or LoadCredential=; recorded as a stage 8 item so
it is a known gap rather than an assumed guarantee.

Verified each identity can do its own job and not the other's:

              tireless   tireless-poller
  add label      403          200
  fork repo      works        403
  read issues    200          200

One measurement caveat recorded in §6.4: the poller can still read repository
contents despite Code=none, because lair/tireless is public. The unit permission
bites on private repos — do not read that as the grant being wrong.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013TxK1CWPkFXqdcXMJ4hVe6
2026-08-07 18:31:59 +03:00
rob thijssen
8586da2cb9 fix(infra): grant the runner's account nothing; split labelling off
All checks were successful
deploy / deploy (push) Successful in 5m34s
Reverses the org team added an hour ago. It was solving the wrong problem.

Measured on the bare account, with no team and no collaborator role: pushing to
its own fork works, opening a cross-repo pull request works, commenting works,
creating an issue works. Pushing to a lair repo is denied and so is deleting it.
Exactly one §2.2 capability is out of reach — applying a label — and that is the
one that must not be bought by granting code write.

So the runner gets nothing at all, which is a better property than a carefully
scoped grant: there is no permission to review, no team to audit, and no rule to
forget on the next repo added. Adding a repo to its remit is now one step, a
fork, rather than three.

Labelling moves to a separate identity, which lines up with the split the two
units already have. The runner records state in Postgres — already the authority
per §2.2 — and the poller reconciles labels onto the forge under its own account
with issue write and no code access. The token that can touch issues is held by
the process that never runs an agent; the token held by the process running
unattended agents can only push to a repository nobody depends on.

That also removes a silent failure I found while testing this. Creating an issue
*with* labels as an unprivileged user returns 201 and drops the labels — no
error, an issue that never gets picked up. A Plan job stamping inherited
admission on its children would have failed precisely that way and looked fine.
Children are now created bare, with their intended labels recorded in Postgres
for the poller to apply.

The second identity is provisioned with stage 2, when the first label is
written. Nothing before then needs it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013TxK1CWPkFXqdcXMJ4hVe6
2026-08-07 18:26:01 +03:00
rob thijssen
b1e01777ff feat(infra): provision the bot account, and move to a fork-based PR model
All checks were successful
deploy / deploy (push) Successful in 5m36s
The design had tireless pushing `tireless/*` branches into the repo it was
working on, with branch protection denying it `main`. It now works the way an
outside contributor does: pushes only to its own fork, reaches the real
repository through a pull request, and has no write access to that repository at
all.

That is a stronger guarantee than the one it replaces. Branch protection is a
rule that can be edited, applies per repo, and is easy to forget on the next repo
added; a bot with no push permission cannot write to `main` whatever anyone
forgets.

The wrinkle is the label protocol. Labels are a write on the issues unit, and
Gitea collaborator permissions are repo-wide, so granting issue write as a
collaborator would hand back the code push the fork exists to avoid. Access is
therefore an org team with unit-level permissions -- lair/tireless-agent, with
Issues=write, PullRequests=write, Code=read, everything else none, scoped to
named repositories rather than the whole org.

Provisioned and verified end to end rather than assumed:

  push to the fork            succeeds
  push to lair/tireless       "User permission denied for writing."
  label an issue              200 / 204
  cross-repo PR from fork     opened as `tireless`
  delete the repository       403

Two credentials with different blast radii: an ssh key on the account for git
transport, and an API token scoped write:issue + write:repository + read:user for
issues and PRs. known_hosts is pre-seeded, because an unattended git must not
prompt and accept-new would trust whatever answered first.

Stage 4 gains the consequence: two remotes, and a fork that has to be checked for
staleness before branching.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013TxK1CWPkFXqdcXMJ4hVe6
2026-08-07 18:18:31 +03:00
rob thijssen
ae285f5597 docs: stage 0 complete — agent login done, all three units active
All checks were successful
deploy / deploy (push) Successful in 5m35s
The interactive login is the one step in stage 0 that cannot be automated, and
it is now done. preflight reports a present subscription login; tireless-runner
is active alongside the api and poller.

Recording it because the previous text said the runner was in `failed`, which
was true for about an hour and is now exactly the kind of stale status this
repo keeps trying not to leave lying around.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013TxK1CWPkFXqdcXMJ4hVe6
2026-08-07 17:59:51 +03:00
rob thijssen
bd658392ab docs(infra): fix the agent login command — cd before npx
The documented command fails:

  npm error Error: spawn sh EACCES
  npm error path: '/home/grenade'

`-H` sets HOME, but sudo leaves the working directory where it was invoked —
the operator's own 0700 home, which the service account cannot read. npx then
fails spawning its `sh -c claude` there. The error names npm and a package, so
it reads like a broken install rather than a directory permission, which is a
bad first experience of the one step in stage 0 that has to be done by hand.

Add the cd, use `bash -c` rather than `sudo -iu` (the account's shell is nologin
by design), explain why both are needed, and point at CLAUDE_PACKAGE so the
version logged in with cannot drift from the version the runner npx's.

Verified on bob: the pinned package now runs as the service account and reports
2.1.220.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013TxK1CWPkFXqdcXMJ4hVe6
2026-08-07 17:54:59 +03:00
rob thijssen
ffa2ad7f72 docs: record stage 0 as deployed, and what the first deploy cost
All checks were successful
deploy / deploy (push) Successful in 5m32s
The status said "nothing has been deployed to bob, the workflow has never run",
which stopped being true at run 8. Leaving it would be the same class of
misleading state this repo keeps trying to remove.

Also record the seven faults it took to get green, because six of the eight were
silent rather than loud: a runner label that meant the job was never scheduled,
a vhost that nginx -t accepts and the SNI router never reaches, a cert whose SAN
only the client checks, an ordering trap, a --chmod that stops applying after
the first deploy, an API that looks healthy from the host it is unreachable on,
and a health probe checking a unit name that expanded to nothing.

The pattern is the useful part for later stages: the expensive faults were the
ones where a check passed while measuring nothing.

Closes #9

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013TxK1CWPkFXqdcXMJ4hVe6
2026-08-07 17:06:59 +03:00
rob thijssen
8067cde288 fix(deploy): expand $unit locally in the health probe
All checks were successful
deploy / deploy (push) Successful in 5m39s
Run 7 shipped everything, brought api and poller up, served /v1/ready through
the proxy — and failed on the last step with exit 4.

  ssh gitea_ci@"$API_HOST" "systemctl is-active \$unit.service"

The escaped \$unit is a leftover from when this ran inside a heredoc. Over a
plain ssh it sends a literal $unit for the *remote* shell to expand, against a
variable that only exists in the workflow's shell — so it ran
`systemctl is-active .service`, which is exit 4, unit not found. It was checking
nothing, and would have kept reporting healthy no matter what the units did.

Also pass -n to both probe sshs. ssh reads stdin by default, and that is a trap
waiting for whoever next pipes something into this step: the first ssh eats the
rest of the script and the remaining checks silently never run. (It bit the test
harness for this fix, which is how it was noticed.)

And say in a comment why tireless-runner is not probed: it refuses to start
until the interactive agent login exists, so probing it would make every deploy
red for something the deploy cannot fix.

Verified by extracting this step's body from the YAML and running it as a file,
the way Actions does:

  {"config":"ok","database":"not_implemented","forge":"not_implemented"}
  tireless-api: active
  tireless-poller: active

Refs #9

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013TxK1CWPkFXqdcXMJ4hVe6
2026-08-07 16:59:49 +03:00
rob thijssen
42582cb922 fix(api): bind the address from config, not a clap default
Some checks failed
deploy / deploy (push) Failing after 5m45s
Run 6 started every unit cleanly and still failed, at the health probe. The
journal says why:

  "tireless-api listening", addr: "127.0.0.1:23296"

while the config it had just loaded says `bind = "0.0.0.0:23296"`. The `--bind`
flag carried `default_value_t`, so it was always Some, so it always won — the
config field I added was never read by anything.

The failure mode is the interesting part: from bob the service looks perfect. It
starts, logs "listening", and answers a local curl. Only a request from the
proxy that actually fronts it fails, which is why the probe was moved off
loopback in the first place.

Make --bind an Option with no default and fall back to config.api.bind, and add
a test asserting the shipped template does not bind loopback — with the reason,
so that if ingress ever moves onto bob the test explains that the bind, the
firewalld service and the vhost move together.

Verified end to end on the real hosts: the API binds 0.0.0.0, the proxy reaches
/v1/ready across the mesh, https://tireless.internal serves both the dashboard
and the API, and the served certificate matches the one on disk by serial.

Refs #9

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013TxK1CWPkFXqdcXMJ4hVe6
2026-08-07 16:52:09 +03:00
rob thijssen
a5efa1e096 fix(deploy): make config readable by the service account, and assert modes
Some checks failed
deploy / deploy (push) Failing after 7m11s
Run 5 shipped everything and started all three units, which then failed
identically:

  configuration is invalid: /etc/tireless/config.toml: Permission denied

The config went out 0640 root:root, but the services run as `tireless`, which
is not in the root group. `--chown root:tireless` is not available as a fix:
on a fresh host that group does not exist until systemd-sysusers runs later in
the same deploy — the same ordering trap that produced the StateDirectory
change. So 0644 root:root, which is defensible precisely because this file
carries no secrets by design; the one that does, tireless.env, stays
0640 root:tireless and is installed by hand.

Testing that on the host turned up a second, quieter fault: `--chmod` only
applies to files rsync actually transfers, so a redeploy whose config is
byte-identical leaves the old mode in place. A mode fix would have appeared to
work on the deploy that introduced it and silently not applied afterwards. Add
-p to the shared options and switch to --chmod=F644/F755 so every push asserts
the mode it wants rather than hoping the content changed.

Also give tireless-runner an explicit start limit. Its preflight fails
permanently without an agent login, and Restart=on-failure with RestartSec=30s
never trips the default 5-starts-per-10s limiter, so it would have retried
forever. Three attempts in ten minutes, then failed, where systemctl status
says why.

Verified on bob: api and poller active, /v1/ready returns
{"config":"ok","database":"not_implemented","forge":"not_implemented"}, and
`tireless preflight` reports Subscription billing with the OpenCode lane
asserted non-Anthropic.

Refs #9

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013TxK1CWPkFXqdcXMJ4hVe6
2026-08-07 16:41:12 +03:00
rob thijssen
267cb3314d fix(deploy): quote --rsync-path, and let systemd own the state directory
Some checks failed
deploy / deploy (push) Failing after 5m49s
Run 4 failed at `ship artifacts` with `sudo: unrecognized option '--server'`.

`R="--rsync-path=sudo rsync --mkpath"` expanded unquoted as `rsync $R` splits
into three arguments — `--rsync-path=sudo`, plus a stray `rsync` that rsync
reads as a source path — so the remote end ran `sudo --server`. Use an array.
The dashboard step quoted it inline and was unaffected, which is why only half
the deploy was broken.

Every rsync destination and every sudo command in `apply system state` has now
been exercised directly against bob as gitea_ci, rather than by another six
minute round trip: seven rsync targets, sysusers, restorecon, firewalld and
daemon-reload all pass.

That surfaced the second fault. restorecon was given /var/lib/tireless, which
does not exist on a fresh host: infra-setup.sh tried to create it before
systemd-sysusers had created the account to own it, so the attempt always raced
and always lost. Declare StateDirectory=tireless on all three units instead —
systemd creates the directory, owns it as the service user and labels it — and
drop the path from restorecon, the grant, and infra-setup.

Refs #9

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013TxK1CWPkFXqdcXMJ4hVe6
2026-08-07 16:32:31 +03:00
rob thijssen
2e3aff7b17 fix(infra): drop the proxy web root from bob's restorecon grant
sudo matches an argument vector exactly, so the sudoers line and the command
the workflow runs have to agree character for character. Moving the dashboard
to the proxy shortened the workflow's restorecon to five paths but left the
grant listing six, which denies it.

/var/www/tireless was never bob's to relabel in any case — the proxy's own
grant covers it, and that one is scoped to static files alone.

Refs #9

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013TxK1CWPkFXqdcXMJ4hVe6
2026-08-07 16:25:27 +03:00
rob thijssen
3c7d95edf9 fix(deploy): correct the runner label, the vhost listen line and the cert paths
Some checks failed
deploy / deploy (push) Failing after 5m31s
Three faults, any one of which would have failed the first deploy — found by
reading the house conventions in ~/git/architecture rather than by running it.

`runs-on: fedora-43-rust` is not a registered runner label. The catalogue in
gitea-runners.md §2 has `rust`; the job would never have been scheduled. Also
collapse build and deploy into one job: the `rust` image descends from
runner-fedora-44 and carries node, ssh and rsync, so the artifact upload and
download — on actions/*-artifact@v3, EOL upstream — bought nothing but a
round-trip and a directory-structure assumption. rpm.lair.cafe's working deploy
is single-job for the same reason.

The nginx vhost bound `listen 443 ssl`. TCP 443 on hanzalova belongs to the
stream SNI router (reverse-proxies.md §4-5, and the live bench.internal.conf
confirms it), so the vhost would never have been reached — the router answers
first with whichever certificate its default branch holds. `nginx -t` passes
either way, which is what makes this worth catching by reading rather than by
deploying. Now `listen 127.0.0.1:14443 ssl proxy_protocol;`.

Cert paths pointed at the host identity cert under /etc/pki/tls. A per-service
name needs its own cert (internal-tls.md §2): the host cert's SAN is bob's FQDN,
not tireless.internal, so verification would have failed. Now
/etc/nginx/tls/{cert,key}/tireless.internal.pem, minted with --san and renewed
by step@tireless.timer.

infra-setup.sh now provisions the ingress rather than describing it: mints the
cert through the JWK provisioner (removing the credential even on failure),
installs the vhost via sites-available + symlink, and registers the
split-horizon record on BOTH routers — a record on one router NXDOMAINs at the
other site. opn-cli has no reconfigure verb, so the apply is a direct API POST;
without it the name resolves only whenever Unbound next happens to reload.

Also fold the stale-bindings check into the workflow (closes the CI half of #8)
and let the runner unit fail without failing the deploy, since it correctly
refuses to start until the interactive login exists.

Refs #9

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013TxK1CWPkFXqdcXMJ4hVe6
2026-08-07 16:23:32 +03:00
rob thijssen
835e3e98f2 fix(deploy): put ingress on the proxy, and make the lint script runnable
Three artefacts disagreed about where nginx runs. design.md §6.2 and the vhost
both said the hanzalova proxy; the API bound 127.0.0.1 and the workflow rsynced
the dashboard to bob. That combination deploys green and then serves nothing,
since a proxy on another host cannot reach bob's loopback.

Resolve it the way design.md already stated: nginx on the proxy, dashboard
shipped there, API bound 0.0.0.0 behind firewalld and the mesh. The health probe
now runs from the proxy over the mesh rather than from bob's loopback, so it
fails when firewalld is closed instead of passing regardless. infra-setup.sh
grows a proxy grant scoped to static files alone, and the nginx vhost install as
a manual step — it needs a certificate, and nothing was telling the operator to
install it at all.

npm run lint had never run: eslint 9 needs a flat config and there was none. Add
it, ignoring the ts-rs generated bindings, and run it in CI so it stays true.

Untrack dashboard/tsconfig.tsbuildinfo, a build artifact that would have put a
spurious diff in every pull request tireless opens.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013TxK1CWPkFXqdcXMJ4hVe6
2026-08-07 15:36:54 +03:00
rob thijssen
ddb2574bbe docs: state the purpose, the autonomy boundary and the dogfooding plan
design.md described what tireless does to an issue but never what it is for, and
in one place said the opposite of the intent: §1 assigned "identifying what
should be worked on" to the operator, which is the thing discovery automates.
A reader — human or agent — would have concluded the system is human-triggered
only and never built the lane.

Rewrite §1 around continuous multi-repo development. Add §2.5, the autonomy
boundary, as a table of which transitions are automatic and which are human,
with the alternatives considered and why they lose. Add §2.6 on anchoring
discovery to a tracking issue rather than making Job.issue a sum type, and what
that buys. Add stage 6 for the lane, renumbering scheduling to 7 and hardening
to 8, with the reasoning for placing a cheap stage late.

Add §10 on tireless working on tireless: that a merged PR restarts the runner
that opened it and why that is survivable but not free, which areas must be
routed to the stronger lane because they are constraint-bearing, and the four
questions dogfooding is expected to answer.

Two new invariants. Admission is inherited (§2.5), and every constraint must be
reachable from a binary's startup path — the rule the previous commit's orphaned
guards violated.

AGENTS.md symlinks CLAUDE.md: both agents look for their own filename, the
implementation prompts tell them to, and a symlink is the only version of this
that cannot drift.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013TxK1CWPkFXqdcXMJ4hVe6
2026-08-07 15:36:54 +03:00
rob thijssen
581e6ae738 feat(discover): add the discovery lane and the autonomy boundary
The docs described a reactive executor: every entry point was a human label, and
the only issue tireless ever created was a plan child. Nothing surveyed a repo or
proposed work, which is the half that makes this continuous rather than
on-demand.

Add JobKind::Discover, routed always to Claude Code (proposing work is the
highest-judgement, lowest-volume task), the tireless/discover and
tireless/proposed labels, and prompt/discover.cc.md as a fourth member of the
versioned prompt set. The contract version does not move: the plan structure is
unchanged, and bumping for less than a shape change trains people to bump
reflexively.

With discovery comes the question of where the loop closes, which was previously
unspecified — routing inferred that plan children are auto-admitted, but nothing
said so. State it as a rule and enforce it:

  Admission is inherited, never invented.

may_opt_in() lets tireless label a plan child, because a human admitted its
parent, and refuses to label a discovered issue, because nothing has been
admitted. It is a function rather than a config flag on purpose: the failure it
prevents is unbounded, not merely wrong, so relaxing it should require review.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013TxK1CWPkFXqdcXMJ4hVe6
2026-08-07 15:36:35 +03:00
rob thijssen
c49b531bcd feat(config): load and validate configuration in every binary
The invariants in CLAUDE.md were tested functions nobody called. There was no
config type at all: --config was accepted and ignored by all three binaries,
figment was an unused dependency, and the 124-line config.toml.tmpl was
aspirational. assert_not_anthropic could not run at startup because nothing read
the provider it checks.

Add tireless_core::config, validated on construction so a Config in hand is a
checked one, and tireless_agent::preflight, called by the worker at startup and
by `tireless preflight` on demand — the same function, so the two cannot
disagree. Move the Anthropic guard to tireless_core::policy where validate() can
reach it; tireless_agent::opencode re-exports it so the documented path resolves.

Also make /v1/ready honest. It returned "ok" unconditionally while its own doc
comment promised dependency checks, so the deploy probe greened on a process
that could do nothing. It now reports per-dependency state, with unwired ones
saying not_implemented rather than ok.

Tests parse the shipped template rather than a fixture, so template and code
cannot drift apart silently.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013TxK1CWPkFXqdcXMJ4hVe6
2026-08-07 15:36:25 +03:00
rob thijssen
21c35ff2c3 docs(prompt): record helexa#179 as settled; shift the oc risk to precedence
Some checks failed
deploy / build (push) Has been cancelled
deploy / deploy (push) Has been cancelled
Passthrough is verified and pinned by regression tests that assert on what
cortex forwarded upstream, so tireless can rely on it. The design no longer
carries it as an open dependency.

Two residual behaviours replace it, neither a passthrough defect:

Multiple system messages are forwarded unmerged and in order, last one wins.
OpenCode sends its own preamble alongside an agent's configured prompt, so the
stage 5 question becomes "did implement.oc.md arrive last", not "did it arrive".
That failure is silent -- an agent that reads as a generic coding assistant and
ignores Out of scope -- and would present as a prompt-quality problem rather
than a plumbing one, so stage 5 now asserts on the upstream request.

helexa#223: /no_think is ignored on /v1/responses, where a thinking model can
spend its whole output budget on reasoning and return "" with status
incomplete. Added RunOutcome::OutputBudgetExhausted so that is retryable but
blameless -- a mis-sized ceiling must not trip a circuit breaker on a lane with
nothing wrong with it. Config pins the chat/completions surface.

The qwen3_next gap is recorded as a gap, not a hole: the system slot is not
arch-branched, so there is no family-specific path to fail.

Fleet now offers Qwen3-Coder-Next, a better fit for executing a written spec,
but cold and feasible only on beast where the pinned 27B lives. Recorded as an
operator decision per generic.md §14 rather than taken here. Config also notes
why tireless pins a model name and never a helexa/* capability alias.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DHhHtohxcdk1PL3tfnYJdH
2026-08-02 15:50:57 +03:00
rob thijssen
7b8308d34e feat(prompt): make the plan handoff a versioned, validated contract
Some checks failed
deploy / build (push) Has been cancelled
deploy / deploy (push) Has been cancelled
helexa#179's application-owned system prompts let tireless shape both ends of
the cc->oc handoff, so the "will a 27B execute an Opus plan" risk becomes a
tunable rather than a hope.

Three prompts, versioned as one set: plan.cc.md tells Claude Code it is writing
for a literal, absent reader; implement.oc.md tells OpenCode to execute exactly
that and report rather than improvise; implement.cc.md covers unplanned issues.
PromptSet::load refuses a mismatched set, and tests assert the prompts mention
every section ChildSpec requires.

The middle is validated, not trusted. plan::validate rejects a plan before any
implementation job is enqueued unless every child carries a runnable acceptance
command (a stopping condition) and a non-empty out-of-scope list (a boundary) --
the two sections a small model needs and a human reader does not. Dangling and
cyclic dependencies are caught too, and implementation_order derives the start
order.

cc uses --append-system-prompt, never --system-prompt: replacing Claude Code's
default discards the tool-use scaffolding that makes it a coding agent. Pin
bumped to 2.1.220, the version this flag surface was verified against.

The oc path depends on helexa#179's passthrough guarantee, which is still open
and unverified for qwen3 arch templating. Stage 5 now opens with a PONG probe
rather than debugging it through a failed implementation run.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DHhHtohxcdk1PL3tfnYJdH
2026-08-02 15:26:05 +03:00
rob thijssen
4e42f87576 feat(tireless): scaffold workspace, dashboard and staged design plan
Some checks failed
deploy / build (push) Has been cancelled
deploy / deploy (push) Has been cancelled
Autonomous issue-to-PR driver for Claude Code and OpenCode, structured per
lair/architecture generic.md.

Workspace: entities/core/data/agent library crates plus api, worker and cli
binaries. Two pieces of real logic land with tests — lane routing (cc for
judgement, oc for specification) and the limit governor.

Constraints encoded as code rather than comments:
- agents are spawned as vendor binaries; tireless never calls a provider API
- ANTHROPIC_API_KEY is never set by tireless, only passed through
- assert_not_anthropic refuses to start an OpenCode lane pointed at Anthropic
- every run passes the governor; provider rate-limit signals win over our own
  accounting

Deployment assets target bob.hanzalova.internal:23296 (registered in
port-allocations.md), fronted by hanzalova at tireless.internal.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DHhHtohxcdk1PL3tfnYJdH
2026-08-02 12:46:42 +03:00