refactor(neuron): cut mistralrs/llamacpp, scaffold candle harness

Stage 1 of the candle-native pivot. Replaces the external-process harness model (mistralrs over HTTP, llamacpp placeholder) with an in-process Harness trait whose sole implementation is candle. The trait keeps its shape so future engines slot in additively, but start/stop default to no-ops and HarnessConfig drops endpoint and systemd_unit since no harness needs external supervision. Behaviour is unchanged on the wire: load_model returns a "not implemented yet (Stage 2)" error and list_models is empty. The gateway-side proxy, poller, and router are untouched. CLAUDE.md Phase 11 (llama.cpp) and Phase 12 (mistral.rs COPR) are marked superseded; the staged plan lives in ~/.claude/plans/create-a-more-aggressive-calm-naur.md. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-18 15:53:04 +03:00
parent 7f797b0265
commit 3cccc2c56b
19 changed files with 203 additions and 401 deletions
--- a/helexa-neuron.spec
+++ b/helexa-neuron.spec
@@ -37,8 +37,9 @@ Provides:       user(neuron)

 %description
 Neuron is a per-node daemon for cortex inference clusters. It discovers
-local GPU hardware via nvidia-smi, manages inference harnesses (mistral.rs,
-llama.cpp), and exposes an HTTP API for model lifecycle management.
+local GPU hardware via nvidia-smi, runs in-process inference via
+huggingface/candle, and exposes an HTTP API for model lifecycle
+management (load, unload, list, inference endpoint).

 %prep
 %autosetup