Add incident grouping (roadmap v0.8: incident list/show/export)

Collapse related alerts into one incident, process-lineage first with a
time-window fallback. Each alert batch's flagged PIDs are walked up the
/proc PPid chain into a lineage set (excluding pid 0/1 so everything
doesn't correlate through init); batches whose lineage sets intersect —
sharing a process or a common ancestor like the web server or SSH
session — join the same incident. PID-less batches (FIM drift, package
tamper, hidden modules) fall back to the most recently active open
incident within incident_window.

snapshot.capture now computes lineage and records each snapshot into a
JSON incident index (incidents.json), writing incident_id into both the
report JSON (report level, so the per-alert schema is untouched) and the
text header. New commands:

  incident list              incidents newest-first, with signatures
  incident show <id>         summary + time-ordered snapshot timeline
  incident export <id>       JSON bundle: record + inlined snapshots

The lineage/assign cores are pure functions; record() serializes the
index under a lock (capture runs from sweep + eBPF threads) and is
best-effort so an index problem never loses a snapshot. 17 new tests
(lineage, assign, record, and an end-to-end capture→group→CLI path).
Config: incident_tracking / incident_window / incident_lineage_depth.
Docs + sample config updated; closes the last v0.8 roadmap item.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
Luna 2026-06-10 17:52:42 -07:00
parent 4015ec872b
commit a56d72edd6
9 changed files with 545 additions and 13 deletions

View file

@ -92,7 +92,7 @@ When a fresh alert survives cooldown, Sentinel writes:
| Interface | Role |
|---|---|
| CLI | Run daemon, one-shot checks, baseline management, FIM checks, package DB checks, rootcheck, posture audit, triage, watchdog. |
| CLI | Run daemon, one-shot checks, baseline management, FIM checks, package DB checks, rootcheck, posture audit, incident review, triage, watchdog. |
| Logs | Durable local evidence under `/var/log/enodia-sentinel`. |
| Dashboard | Read-only web view for status, alert browsing, and snapshot inspection. |
| Push | ntfy, Pushover, and generic webhook notifications. |
@ -164,8 +164,11 @@ The current `Alert` object is the public detection unit:
| `detail` | Human-readable explanation. |
| `pids` | Process IDs for snapshot deep dives. |
Future incident grouping should add an `incident_id` around one or more alerts,
without breaking the alert JSON schema.
Incident grouping adds an `incident_id` at the snapshot-report level (not inside
each `Alert`), so it groups one or more alerts without changing the alert JSON
schema. Grouping is process-lineage first, with a time-window fallback for
PID-less alerts; see the [command reference](COMMAND_REFERENCE.md) and
[roadmap](ROADMAP.md).
## Configuration Model