Vetting

Author	SHA1	Message	Date
josh	017c3c38fe	feat(ui): 15-point UX overhaul — affordances, feedback, and navigation CI / Lint + build + test (push) Successful in 1m43s Details Release / detect (push) Successful in 6s Details Release / build-live-image (push) Has been skipped Details Release / bundle (push) Successful in 52s Details Address friction points identified in a full interface audit: - Re-add status badge to dashboard tiles so run state is visible at a glance - Add active nav indicator and SSE connection health monitor (live/stale) - Show manual registration form by default instead of hiding behind <details> - Add copy-to-clipboard buttons on SSH hold command and quick-register one-liner - Replace tooltip-only profile descriptions with inline visible text - Clarify non-destructive toggle with explicit stage impact description - Replace disabled "Start vetting" button with actionable offline guidance - Swap browser confirm() dialogs for styled inline confirmations - Add colored badge to spec diffs summary visible when collapsed - Add distinct "cancelled" mood for cancelled runs (vs idle) - Add match count to log search and aria-label for accessibility - Add styled 404 page rendered inside the app shell Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>	2026-04-23 20:08:07 -04:00
josh	17ec55cb85	chore: cleanup sprint — dead CSS, dedup helpers, handler refactor CI / Lint + build + test (push) Successful in 1m34s Details Release / detect (push) Successful in 4s Details Release / build-live-image (push) Has been skipped Details Release / bundle (push) Successful in 1m5s Details Remove ~126 lines of orphaned CSS from tile slim-down and old detail layout. Consolidate 4 duplicate duration formatters into shared elapsed()/fmtElapsed() helpers. Break 160-line Result handler into focused sub-functions. Implement real Hub.Shutdown() (was a no-op). Standardize agent error responses to JSON. Replace panic() in router init with error return. Extract magic numbers as named constants. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>	2026-04-21 20:39:38 -04:00
josh	c11573eeeb	feat(ui): slim dashboard tile to hostname + online/offline only CI / Lint + build + test (push) Successful in 1m33s Details Release / detect (push) Successful in 5s Details Release / build-live-image (push) Has been skipped Details Release / bundle (push) Successful in 53s Details Run status, Start/Cancel/View controls, and non-destructive toggle all live on /hosts/{id} — duplicating them on the dashboard tile clogged the grid and wouldn't scale past a handful of hosts. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>	2026-04-20 22:56:05 -04:00
josh	e73b221a8c	fix(ui): fit pipeline timeline without horizontal scroll CI / Lint + build + test (push) Successful in 1m39s Details Release / release (push) Successful in 7m30s Details 15 nodes (3 pre-stage + 11 stage + Completed) exceeded the 1280px main container's usable width, producing a horizontal scrollbar under the pipeline on the run page. Widen main to 1440px, tighten per-node min widths, drop the scrollbar, and split camelCase labels so multi-word stages ("WaitingReboot", "SpecValidate", "CPUStress") wrap onto two lines instead of forcing node width. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>	2026-04-19 22:51:10 -04:00
josh	23c689aa5b	deep profile + threshold gating + firmware stage + Burn super-stage CI / Lint + build + test (push) Failing after 1m57s Details Release / release (push) Has been cancelled Details Ships all five phases of the deep-profile overhaul together. Runs now carry a profile (quick/deep/soak); every profile walks the same 11-stage order — Inventory → Firmware → SpecValidate → SMART → CPUStress → Storage → Network → Burn → GPU → PSU → Reporting — with only per-stage durations and concurrency scaled. Phase 1: profiles.ProfileRegistry loaded from vetting.yaml; runs.profile column + CreateWithProfile; threshold table + evaluator seeded per-run from the shared vetting.thresholds block; breach flips result at /sensor + /result. Phase 2: upgraded CPUStress (stress-ng --cpu-method=all --verify + EDAC/MCE poll), Storage (fio --verify=md5 + SMART start/end delta), Network (sustained iperf + /proc/net/dev deltas) with per-profile knobs from Deps. Phase 3: Burn super-stage with goroutine fan-out for CPU + memory + fio + iperf, PSU rails sampled across the Burn window, SensorMux (2 s flush, 500-sample cap) to absorb backpressure. Phase 4: Firmware stage + firmware_snapshots table; probes dmidecode (BIOS), ipmitool (BMC), ethtool -i (NIC), nvme (sysfs + id-ctrl), lspci (HBA), /proc/cpuinfo (microcode). spec.DiffFirmware folds into SpecValidate with pin-by-identifier and fan-out-across-component matching; mismatches park the run in FailedHolding. Phase 5: profile radio on the host start form, profile chip on the run header, Firmware section in the HTML report, coverage artifact uploaded from CI, agent/tests/fakes/ scaffold with Deps.LookPath seam + stress_ng and dmidecode example fakes. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>	2026-04-18 22:50:57 -04:00
josh	19608bef1b	ui: split /hosts/{id} into host page + /runs/{runID} run page CI / Lint + build + test (push) Successful in 1m35s Details Release / release (push) Successful in 23m47s Details Host page owns host metadata, full runs table with per-row stage strip, in-flight banner, and empty-state CTA. Run page owns pipeline, active step, logs, sub-steps, spec diffs, and hold banner with a breadcrumb back to the host. Dashboard tile reverts to host-only. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>	2026-04-18 20:37:57 -04:00
josh	5c6bfa5ffa	ui: fix log lines rendering vertically when stage prefix is present CI / Lint + build + test (push) Successful in 1m39s Details Release / release (push) Has been cancelled Details The .log-line grid was templated with 5 columns (anchor/ln/lvl/ts/text), but renderLogSSE inserts an optional log-stage span, making 6 children. The 6th child wrapped to row 2 column 1 (24px wide), which forced the message text to break one character per line. Flexbox with min-width:0 on the text span scales cleanly with or without the stage element. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>	2026-04-18 19:20:51 -04:00
josh	f79fe0f0db	ui: GitHub-Actions-style detail page, sub-steps, mini-tile run-view CI / Lint + build + test (push) Successful in 1m26s Details Release / release (push) Successful in 6m47s Details Reshapes the detail page into a run-view: hybrid horizontal pipeline + expanded active-step pane with sub-steps, a per-step log pane with line-numbered permalinks and client-side search, and a runs-history sidebar that navigates via ?run=N. Default step is server-picked (running → failed → Reporting) so the operator lands on the thing that's moving. Adds a sub_steps table + SSE topic (substep-{run}-{stage}-{ordinal}) so per-disk and per-pass work (SMART, CPUStress CPU/RAM, Storage, GPU) is visible in the UI instead of buried in stage summary JSON. Agent emits sub-step reports from existing per-iteration loops. Dashboard tiles become a mini run-view with a 9-dot step strip so the operator reads run health across the whole grid at a glance. Register page gets the same card shell + button styling. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>	2026-04-18 19:00:11 -04:00
josh	1694c20b12	Host detail v2: full pipeline + per-stage logs + WoL diagnostics CI / Lint + build + test (push) Has been cancelled Details Pipeline now always renders all 13 nodes (3 pre-stage + 9 stage + Completed), synthesising ghosts from run state when stage rows aren't seeded yet. Makes a WaitingWoL host show the full timeline ahead of it instead of just 4 dots. Agent tags each log line with its stage; logs.Hub fans out to both log-{runID} and log-{runID}-{stage} SSE events so the detail page can show per-stage tabs with a pure-CSS radio-sibling switch. Flat run log prepends [stage] so grep still works. Dispatcher writes picked/sent-WoL/heartbeat lines into the per-run log — the operator opens the detail page, sees WaitingWoL stuck, and reads exactly what the dispatcher did and why nothing's progressing, instead of having to tail journalctl on the LXC. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>	2026-04-18 00:38:27 -04:00
josh	bb658a8435	Host detail page + pipeline timeline CI / Lint + build + test (push) Has been cancelled Details Click a tile to open /hosts/{id} — the canonical control surface per host. Timeline renders every pre-stage, stage, and terminal node in order, with the current one pulsing, failed ones flagged, and downstream ones dimmed as skipped. Detail page shows summary, hold card (when holding), all action buttons, spec diffs, a full-height log pane, and a collapsed expected-spec YAML. Tile slims to name, last-seen, status, and one primary action; a CSS-overlay <a> makes the whole card clickable while buttons stay receptive via z-index. Runner.publishTileUpdate now also emits pipeline-{runID} fragments, and CompleteStage wraps Stages.CompleteByName so stage completions advance the timeline live — without this the dots only moved on state transitions. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>	2026-04-17 23:59:43 -04:00
josh	a0c0fb114f	Add host-mode heartbeat: vetting-agent host + last-seen badge CI / Lint + build + test (push) Has been cancelled Details vetting-agent gains a `host` subcommand that runs as a systemd service installed by the quick-register one-liner, POSTing every 30s to /api/v1/hosts/{mac}/heartbeat so the dashboard tile shows "online" or "Nm ago" without waiting on WoL. Ships dormant client code for the Phase 2 reboot_for_vetting command so the server can flip it on later without a binary redeploy. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>	2026-04-17 23:34:15 -04:00
josh	8b3d9a312e	Add quick-register one-liner for target-host registration CI / Lint + build + test (push) Failing after 5m15s Details Operator pastes `curl -fsSL $ORCH/register/quick.sh \| sudo bash` on the target host (pre-wipe). The script probes MAC + CPU/RAM/disks/NICs/GPUs, emits an expected-spec YAML, and POSTs to a new LAN-trusted JSON endpoint /api/v1/hosts. The register page shows the command prefilled with the orchestrator URL; the manual form moves into a collapsible "Register manually" disclosure.	2026-04-17 22:50:54 -04:00
josh	9bb4b09a04	Initial commit: full Phases 1-6 implementation CI / Lint + build + test (push) Has been cancelled Details Post-repair hardware validation pipeline for Proxmox cluster hosts. Go orchestrator + in-image agent + mkosi live image + bundled dnsmasq PXE + SQLite + HTMX/SSE UI + notify registry + janitor + full docs.	2026-04-17 21:32:10 -04:00

13 Commits