Architecture
The whole system is four stages in a line: collect, centralize, analyze, alert. Every later page is one of these stages in detail.
The pipeline
- Collect. Every host generates logs already; authentication attempts, firewall decisions, service errors, system messages. The system logger on each host is the source. No new agent is needed because the logging is happening whether you use it or not. The job is to route the useful parts off the host.
- Centralize. Each host forwards its logs to one central collector. From that point on there is a single place that holds a copy of what every host saw, which is the foundation everything else builds on.
- Analyze. On the collector, small scripts read the gathered logs and reduce them to signals: not "here are ten thousand log lines" but "this many failed logins from one address in the last hour," "this host's disk crossed a capacity threshold," "a critical error appeared." Analysis turns volume into meaning.
- Alert. When a script produces a signal worth acting on, an alerting layer delivers it to a human. Most of the time nothing fires; the point is that when something does, you are told rather than having to look.
Why centralize at all
Forwarding logs off each host, rather than analyzing them where they sit, buys three things that matter:
- A compromised host cannot cover its tracks. If an attacker gets onto a host, one of the first things they may do is tamper with local logs to hide what they did. If the relevant events have already been forwarded off the host as they happened, that tampering is defeated. The evidence is already somewhere the attacker does not control. This alone is a strong argument for centralization.
- Correlation across hosts becomes possible. Some things are only visible when you can see several hosts at once, the same source probing multiple services, a pattern that moves from host to host. Logs scattered on their originating machines cannot show you that. Logs in one place can.
- The logs survive the host. If a host falls over, crashes, is compromised, is rebuilt, its local logs may be lost or suspect. Forwarded logs persist independently of the host that produced them.
Where analysis lives
Analysis runs on the central collector, close to the gathered logs, as a set of small single-purpose scripts. Each script owns one category, reads that category's logs for a recent window, and emits a signal if something crosses a line. Keeping each script narrow, one concern each, is what makes the system understandable and maintainable by one person. When an alert fires, you know exactly which small script to look at.
This is deliberately not a big correlation engine. It is a handful of focused checks, which is the right amount of machinery for the job at this scale.
Where alerting lives
The final stage, turning a signal into a notification a human receives, is handled by a scheduling-and-automation layer that runs the analysis scripts on a cadence and delivers their output. This can be as simple as cron and a mail command, or a dedicated automation platform. The alerting page covers both, and the tradeoffs between them, including the fact that the automation layer is itself something you have to secure.
The whole thing in one sentence
Every host forwards its logs to one collector, small scripts on the collector reduce those logs to signals, and an automation layer delivers the signals that matter to a human. That is the entire architecture, and its simplicity is the point. It is built from standard tools, it is understandable in one sitting, and it delivers the core SIEM function without the SIEM product.
No comments to display
No comments to display