Add host correlation and assurance coverage

This commit is contained in:
Luna 2026-07-10 03:07:00 -07:00
parent 8194d13734
commit b40ac4252c
No known key found for this signature in database
36 changed files with 944 additions and 23 deletions

96
docs/FLEET_DESIGN.md Normal file
View file

@ -0,0 +1,96 @@
# Fleet Collector Design
This note sketches the v1.2 fleet layer. It is intentionally optional: a local
Sentinel agent must keep detecting, snapshotting, and serving its local console
even when no collector exists or the collector is down.
## Goals
- Give a small fleet one view of host health, stale sensors, repeated
signatures, incidents, and integrity drift.
- Preserve local autonomy. Detection, alert capture, response-plan generation,
and local evidence retention stay on each host.
- Use append-only, auditable ingest. The collector stores what hosts report; it
does not execute response actions on agents.
- Keep transport simple: HTTPS with pinned/self-signed deployments, or a
tailnet address where the operator already has one.
## Non-Goals
- No remote shell, remote command execution, or automatic containment in the
first fleet version.
- No dependency on the collector for alerting or local dashboard availability.
- No central trust reset. If a host is suspected compromised, the collector can
show evidence and staleness, but recovery remains an operator-reviewed local
workflow.
## Agent-to-Collector Payloads
Agents should push compact JSON records on a cadence and after important local
events:
| Payload | Source | Minimum contents |
|---|---|---|
| Heartbeat | `status --json` | host id, time, version, daemon state, heartbeat age, sensor states |
| Alert summary | alert snapshot | snapshot name, incident id, severity, SIDs, signatures, counts |
| Incident summary | `incidents.json` | id, host, timestamps, severity, SIDs, signatures, correlations |
| Integrity summary | `/api/integrity` | watchdog, FIM/package anchors, reconciliation, hash-chain summary |
| Hash-chain anchor | `hash-chain.jsonl` | latest `enodia.hash_chain.v1` record or last hash |
Large evidence bundles stay local at first. The collector stores paths and
digests so an operator can pull evidence deliberately.
## Enrollment
Each host gets:
- a stable host id,
- a per-host bearer token or mTLS client certificate,
- an optional display name and tags,
- a collector URL.
Enrollment should write local config only after explicit operator approval. A
lost token can be revoked collector-side without changing local detection.
## Ingest API Shape
Initial endpoints can stay narrow:
- `POST /api/v1/ingest/heartbeat`
- `POST /api/v1/ingest/alert`
- `POST /api/v1/ingest/incident`
- `POST /api/v1/ingest/integrity`
- `POST /api/v1/ingest/hash-chain`
Every request includes host id, timestamp, schema id, monotonic sequence number,
and payload digest. The collector rejects unknown hosts, stale timestamps beyond
policy, replayed sequence numbers, and schema ids it cannot parse.
## Collector Storage
A small SQLite database is enough for the first version:
- `hosts`: identity, tags, token/cert fingerprint, last seen, enrollment state.
- `heartbeats`: append-only health samples.
- `alerts`: compact alert summaries keyed by host and snapshot name.
- `incidents`: latest incident summary plus history rows for changes.
- `integrity`: latest integrity state plus historical samples.
- `hash_anchors`: host hash-chain tips over time.
The collector dashboard should answer:
- Which hosts are silent?
- Which hosts have critical incidents?
- Which hosts report repeated SIDs/signatures?
- Which hosts have missing/stale integrity anchors?
- Did a host hash-chain tip change as expected over time?
## Security Notes
- Treat collector data as sensitive forensic metadata; require TLS and auth.
- Store per-host credentials hashed or certificate fingerprints only.
- Do not let the collector mutate agent config in v1.2.
- Keep host-side evidence and response actions local until a separate remote
write-path design is reviewed.
- Consider mirroring `enodia.hash_chain.v1` records off-box before adding any
higher-risk fleet command surface.