Skip to main content

Analysis

Pattern: turning collected logs into signals. Collection gathers raw logs, analysis reduces them to the handful of statements worth acting on. This page shows one category in full, health monitoring, chosen because it is operational rather than a detection tripwire and therefore safe to show in its entirety. Then continues to describe the shape the other categories follow without their sensitive specifics.

The shape of an analysis script

Every analysis script has the same skeleton, regardless of category:

  1. Read the relevant logs for a recent time window.
  2. Reduce them to a small number of facts (a count, a maximum, a change).
  3. Compare those facts against a threshold.
  4. Emit a signal (exit status, output, a log line) if the threshold is crossed.

The alerting layer, covered on the next page, runs these scripts on a schedule and delivers whatever they emit. Keeping each script to this narrow shape, one category, one window, one comparison, is what makes the whole system understandable. When something fires or fails to fire, you know exactly which small script to read.

The worked example: health monitoring

Health monitoring is a good worked example because knowing how it works helps no attacker, disk and CPU thresholds are not a security tripwire, and the pattern it demonstrates is exactly the one every other category follows. It comes in two halves: a collector on each host that emits health snapshots, and a check on the central host that reads them.

Prerequisite: the per-host health snapshot

This half is a prerequisite for the fleet check and worth setting up first. Each host runs a small script every ten minutes that measures its own resource usage and logs a single tagged line. That line is forwarded to the collector like any other log (via the health tag rule shown on the collection page).

#!/bin/bash
# /usr/local/bin/sysstat-health.sh
# Runs on EVERY host, every 10 minutes via cron.
# Emits one tagged syslog line with current resource usage,
# which rsyslog forwards to the central collector.

set -e

HOST=$(hostname -s)
TS=$(date -Is)

# CPU usage as a percentage (user + system), sampled over 1 second.
CPU_USED=$(LC_ALL=C sar -u 1 1 | awk '/Average:/ { printf "%.1f", 100-$8 }')

# Memory usage as a percentage of total.
MEM_USED=$(LC_ALL=C sar -r 1 1 | awk '/Average:/ { printf "%.1f", ($3+$5)/$2*100 }')

# 1-minute load average.
LOAD1=$(awk '{print $1}' /proc/loadavg)

# Root filesystem usage as a percentage.
DISK_ROOT=$(df -P / | awk 'NR==2 {gsub("%","",$5); print $5}')

# Emit one structured line under the "health" tag at info severity.
# rsyslog matches the tag and forwards it to the collector as health.log.
logger -p local5.info -t health \
  "ts=${TS} host=${HOST} cpu=${CPU_USED}% mem=${MEM_USED}% load1=${LOAD1} disk_root=${DISK_ROOT}%"

Scheduled on each host with a cron entry:

# /etc/cron.d/sysstat-health
*/10 * * * * root /usr/local/bin/sysstat-health.sh

The host does not decide whether its own numbers are a problem. It just reports them, every ten minutes, in a structured line. All the judgment lives in one place on the collector, which keeps the per-host piece trivial and means you tune the logic in exactly one location rather than on every host.

The check: reading the fleet's health on the collector

On the collector, a script runs on a schedule and reads the forwarded health snapshots for all hosts over the recent window. Because each host reports every ten minutes, a window of the last hour gives several data points per host, enough to tell a momentary spike from a sustained problem. The script's logic, in shape:

  • For each host, read its health snapshots from the recent window.
  • Determine whether any resource crossed a concern threshold, and whether a transient reading (one spike) should be treated differently from a sustained one (crossing the line repeatedly across the window).
  • Emit a signal naming any host that warrants attention.

The distinction between a single spike and a sustained condition is the useful part of the logic. A host momentarily busy is normal, a host busy across the whole window is a problem, and treating them the same produces either false alarms or missed real issues. Requiring a condition to persist across multiple readings before alerting is what keeps health monitoring quiet enough to trust. The exact thresholds and how many readings count as "sustained" are tuning particular to an environment, and are the kind of specific this book leaves you to set against your own normal.

The other categories, in shape

Every other analysis category, authentication, boundary events, critical severity, service-specific signals, follows the same skeleton as the health check, differing only in which logs it reads and what condition it looks for. Described in general, without the sensitive specifics:

  • Authentication analysis reads the forwarded authentication logs for a recent window, counts failure patterns by source, and emits a signal when a source's behaviour looks like an attack rather than a fumble. The precise counts and the exact patterns that distinguish "attack" from "bad day" are the sensitive tuning, omitted deliberately.
  • Boundary analysis reads the firewall and intrusion-prevention logs, aggregates recent drops and bans, and emits a signal when the volume or pattern from a source suggests a probe or flood rather than background noise.
  • Critical-severity analysis reads the per-host critical logs (which the collection stage already separates out) for the recent window and emits a signal if anything appeared at all, since by definition these are events the systems themselves flagged as serious.
  • Service-specific analysis reads a particular service's logs for the specific signal that service warrants, a new registration, a sensitive action, a known error pattern, and emits accordingly.

In every case the skeleton is identical. Read a window, reduce to facts, compare to a line, emit a signal. Once you have written one, the rest are variations, which is what makes this approach maintainable by one person.

Why scripts rather than a query language

An enterprise SIEM would express these as saved queries or correlation rules in its own language. Small single-purpose shell scripts do the same job here with no product, and they have a real advantage at this scale. They are transparent and debuggable with tools you already know. When a check misbehaves, you read a short script and run it by hand, rather than reverse-engineering a rule engine's behaviour. The tradeoff is that scripts do not scale to thousands of sources or complex cross-source correlation, at which point you would want a real SIEM.