quantus/miner#9. Additive metrics in a lair-owned module of the metrics
crate; origin's metrics are untouched so the fleet dashboard keeps working
across origin merges.
Identity:
- miner_build_info{version, commit}: the join key for everything below.
- miner_config_info{engine, gpu_batch_size, gpu_devices, cpu_workers,
gpu_throttle_ms}: two deploys of one commit with different flags are
different experiments.
Per device (handles resolved once at engine init, no label lookups in the
batch path):
- miner_device_hashes_total{device, kernel}, miner_device_solutions_total,
miner_device_lost_total.
- miner_gpu_batch_seconds{device, kernel, phase=gpu|host}: submit-to-mapped
on the device versus everything else in the batch. This is the bubble
#4, #5 and #6 attack, measured directly.
Jobs and results:
- miner_jobs_received_total, miner_stale_hashes_total{engine},
miner_job_pickup_seconds{engine} (issued to picked up; the cancel latency
of a busy worker), miner_results_submitted_total,
miner_results_send_failed_total, miner_seal_latency_seconds (found to
sent), miner_job_idle_seconds_total (sent to next job, node-attributable).
Connection: miner_connects_total, miner_connect_failures_total,
miner_disconnects_total, miner_connected, miner_disconnected_seconds_total.
Origin-owned code touched, each block marked lair: engine-gpu gains a
metrics dependency, a DeviceMetrics handle on GpuContext and timing points
in run_single_batch; miner-service gains found_at on WorkerResult,
created_at on MiningJob and the recording calls; miner-cli sets build info
at startup. Deploy validate now asserts miner_build_info carries the
deployed commit.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CBgs2nSi4H2mdh8kD8vMX5