enodia-sentinal/docs/ROADMAP.md
2026-06-13 05:53:48 -07:00

11 KiB

Enodia Sentinel Roadmap

This roadmap moves Enodia Sentinel from single-host endpoint detection and response into a fuller host security platform: detect, verify, investigate, respond, and assure. Dates are intentionally not promised here; the sequence matters more than calendar precision.

Release Tracks

Track Goal
Sensor Better host telemetry, event capture, and detection coverage.
Integrity Stronger proof that files, packages, baselines, and Sentinel itself were not silently modified.
Incident Group alerts, build timelines, and export evidence.
Response Add safe, auditable containment and recovery workflows.
Fleet Let multiple hosts report to a small controller without making the local agent dependent on it.
Assurance Move trust anchors off-box or below user space.

v0.8: Documentation, Posture, and Incident Shape

Purpose: make the current agent easier to operate and create the data model for work that is bigger than single alerts.

  • Add incident_id grouping around related alerts (process-lineage first, time-window fallback) while preserving the existing alert JSON — incident_id is added at snapshot-report level.
  • Add an incident timeline generator: incident show builds a time-ordered view across an incident's member snapshots (time, severity, signatures, PIDs). Richer sources (package transactions, dedicated FIM-diff lines) can follow.
  • Add enodia-sentinel incident list/show/export.
  • Add host posture checks (enodia-sentinel posture check): SSH root/password/empty-password login, passwordless and writable sudoers, world-writable PATH components, loose sensitive-file permissions, and disabled package signatures. Still to come: unexpected enabled services and unsafe systemd unit settings.
  • Add a machine-readable status --json command for automation.
  • Update the red-team harness so every current detection has a named drill and expected sid (sentinel-redteam --list), with a PASS/MISS verification pass.
  • Document incident response runbooks for reverse shells, persistence, trojaned binaries, hidden listeners, and sensor tampering (RUNBOOKS.md).
  • Add Gonzalo/Peopleswar-style behavior coverage: input-device keylogging, credential-store/private-key access, SCTP/DCCP/raw/packet-family traffic, and raw ICMP/special-protocol rootcheck detection.
  • Add process-hiding and memory-obfuscation coverage: /proc vs ps cross-view checks plus memory-map scans for hide libraries, RWX memory, and executable anonymous/memfd/deleted mappings.
  • Add event-driven memory telemetry for short-lived mprotect/mmap RWX, memfd_create, ptrace, seccomp, cross-process memory access, and memory locking syscalls.

Exit criteria:

  • One alert can become one incident artifact.
  • Operators can answer "what changed around this alert?" without reading every raw snapshot manually.
  • Posture checks report reviewable findings but do not block daemon startup.

v0.9: Response and Recovery

Purpose: move from "tell me" to "help me act" without unsafe automation.

  • Add response plan generation: enodia-sentinel respond plan <incident-id>.
  • Persist CLI-generated dry-run plans under response-plans/ and append response-audit.log records for review/handoff.
  • Add dry-run first actions: kill process, stop/disable service, block remote IP, quarantine file, restore package-owned file by reinstalling its package, and freeze evidence.
  • Require explicit --apply for changes; default to read-only plans.
  • Extend response audit logs to any future state-changing --apply workflow.
  • Add baseline reconciliation: accept legitimate FIM/package/listener/SUID changes with a recorded reason.
  • Add richer false-positive suppression suggestions that can emit TOML snippets but never modify config automatically.
  • Add recovery checks: "verify package", "re-run rootcheck", "confirm no persistence diff", and "confirm heartbeat/dashboard visible from watchdog".

Exit criteria:

  • An operator can go from alert to reviewed response plan with one command.
  • Any state-changing action is explicit, logged, reversible where possible, and testable in a safe fixture.

v1.0: Stable Local Platform

Purpose: define a stable single-host product.

  • Commit to stable JSON schemas for alerts, incidents, status, and response audits.
  • Add compatibility tests for schema evolution.
  • Add documentation versioning and manpage-style command reference.
  • Harden packaging for Arch first, then document Debian/RPM install paths.
  • Add signed release artifacts and checksums.
  • Improve dashboard from alert browser to local console: incidents, posture, integrity state, response plans, and watchdog status.
  • Keep the dashboard read-only unless a separate authenticated write path is designed and reviewed.

Exit criteria:

  • The local agent can be installed, operated, upgraded, and integrated without reading source code.
  • JSON and command contracts are stable enough for downstream automation.

v1.1: Event Coverage and Correlation

Purpose: reduce blind spots and connect related signals.

  • Generalize the exec-rule engine into a typed host-event rule engine: exec, tcp_connect, bind, listen, accept, file_write, chmod/chown, setuid/setgid, capset, and module-load events.
  • Keep rules declarative and Snort-like: sid, msg, severity, classtype, event, and match fields such as process name, parent process, argv regex, path prefix, peer IP/port, UID/GID, package ownership, and lineage.
  • Add eBPF event sources: tcp_connect, bind, accept, module load, privilege transitions, and writes to watched persistence paths.
  • Build process lineage across polling and event sources.
  • Correlate multi-stage behavior: web service spawns shell, shell downloads payload, payload adds persistence, listener appears, and FIM changes.
  • Add configurable correlation windows and severity escalation rules, including:
    • exec_rule.web-rce + suspicious egress + persistence within 10 minutes → CRITICAL multi-stage intrusion.
    • new_listener + new_suid in the same lineage/window → CRITICAL backdoor with privilege-escalation artifact.
    • fim_modified or pkg_signature_mismatch + pkgdb_tamper → CRITICAL suspected trojaned binary with checksum cover-up.
    • web/database service spawning shell + tool transfer (curl|wget|sh) → CRITICAL likely RCE payload staging.
  • Add high-signal living-off-the-land detections: curl|wget | sh, Python/perl/bash reverse-shell argv variants, systemctl enable from unusual ancestry, chmod +s outside package transactions, interpreter egress to unusual public ports, and execution from writable directories.
  • Add first-seen and rarity tracking for local network behavior: first public destination per process name, first listening port per binary, interpreter connections to non-standard ports, and long-lived low-byte connections when socket sampling can support it.
  • Keep polling as the oracle and fallback.

Exit criteria:

  • Short-lived network and persistence actions are visible.
  • Related alerts collapse into one incident with a readable timeline.
  • Multi-stage correlation raises one higher-confidence incident without hiding the underlying raw alerts.

v1.1.1: Detection Quality and Rule Operations

Purpose: make detection coverage easier to audit, tune, and extend without turning Sentinel into a noisy rules dump.

  • Add enodia-sentinel rules list and rules show <sid> for built-in and configured rules.
  • Add enodia-sentinel rules test <event-json> so operators can validate custom event rules against captured or fixture events.
  • Generate rule documentation from source defaults: SID, signature, classtype, event type, match fields, expected false positives, and drill coverage.
  • Require a fixture or safe red-team drill for every built-in SID, including event-only and correlation SIDs.
  • Add compatibility tests for rule JSON/event schema evolution.
  • Extend triage so false-positive suggestions are tied to SID/classtype and can emit TOML snippets for allowlists, correlation windows, or severity overrides.
  • Add post-alert enrichment consistently across detections: package owner, executable hash, parent/ancestor chain, remote IP classification, file mode/owner, recent writes by same lineage, and current FIM/package/rootcheck status.

Exit criteria:

  • Operators can inspect, test, document, and tune detection rules without reading Python source.
  • Every shipped detection rule has reproducible fixture/drill coverage and stable operator documentation.

v1.2: Fleet Mode

Purpose: support several personal or small-business Linux hosts without turning the agent into a heavy enterprise platform.

  • Add an optional collector service.
  • Agents push signed JSON events and heartbeats over HTTPS or a tailnet.
  • Collector stores host status, alerts, incidents, and package/integrity state.
  • Add host enrollment with per-host tokens.
  • Add fleet-wide views: stale sensors, repeated signatures, hosts missing package signatures, and integrity drift.
  • Preserve local autonomy: detection and evidence capture must continue if the collector is down.

Exit criteria:

  • A small fleet can answer "which host is compromised or silent?" from one console.
  • The agent still works as a standalone local tool.

v1.3+: Assurance and Hardening

Purpose: make tampering harder to hide from an attacker with high privileges.

  • External baseline anchor: mirror FIM and package DB fingerprints off-box with append-only semantics.
  • Hash-chain events.log and snapshots.
  • Optional snapshot signing with a host key.
  • Investigate Linux IMA/EVM and TPM-backed attestation for supported machines.
  • Replace bcc with a libbpf + CO-RE agent if the event layer becomes core to the product.
  • Explore minimal kernel-side enforcement hooks for response actions, while keeping the default product inspect-only.

Exit criteria:

  • A local root attacker has to tamper with both the host and an external or hardware-backed trust anchor to hide compromise.

Backlog

Useful ideas that need design before implementation:

  • Policy-as-code for host posture.
  • YARA-compatible scanning for selected paths.
  • Sigma-like host event rules.
  • SBOM and package provenance reporting.
  • Offline evidence bundle for outside analysis.
  • Dashboard import of exported incident bundles.
  • Support for non-pacman package managers in signed-package verification.
  • Dedicated Debian/RPM packaging.
  • Installer preflight checks and health diagnostics.
  • Rule documentation generated from source defaults.

Roadmap Rules

  • Do not add dependencies to the base daemon without a specific security or operator value that cannot be achieved with stdlib.
  • Prefer read-only detection and dry-run response until the command contract is tested and documented.
  • Every new detector or response action needs a safe drill or fixture.
  • Do not build stealth or self-hiding behavior.
  • Keep the local agent useful without a dashboard, collector, or network.