Skip to main content

Collection

Pattern: getting logs off every host and into one place. This is the plumbing, getting this stage right is what everything else depends on.

The mechanism is the standard Linux system logger (rsyslog), configured on each host to forward, and on one central host to receive. No extra software.

The receiving side: the collector

One host is designated the collector. It runs a listener that accepts forwarded logs from the other hosts and writes them into per-host directories, so each host's logs stay separate and identifiable.

Enable the listener (here on TCP port 514):

# /etc/rsyslog.d/10-listen.conf  (collector only)
module(load="imtcp")
input(type="imtcp" port="514")

Then route incoming logs into a directory per source host. A template builds the path from the sending host's name and the program that produced the log, so logs land at predictable locations:

# /etc/rsyslog.d/40-remote.conf  (collector)

# Critical events (severity <= err) copied to a per-host critical log,
# in addition to their normal destination.
if ($fromhost-ip != "127.0.0.1" and $syslogseverity <= 3) then {
    action(
        type="omfile"
        File="/var/log/remote/%HOSTNAME%/critical.log"
    )
    # no "stop": the message continues to its normal per-program log too
}

# Everything else from a remote host: file it under the host's directory,
# named for the program that produced it.
template(name="RemoteLogs" type="string"
  string="/var/log/remote/%HOSTNAME%/%PROGRAMNAME%.log")

if ($fromhost-ip != "127.0.0.1") then {
  action(type="omfile" DynaFile="RemoteLogs")
  stop
}

The result on the collector is a tidy tree: /var/log/remote/<host>/<program>.log, with a separate critical.log per host catching anything high-severity. That structure is what the analysis scripts read later; predictable paths simplify analysis.

The sending side: every host

Each host forwards the log streams worth centralizing to the collector. A named ruleset defines the forwarding action once, with a queue so that logs are not lost if the collector is briefly unreachable:

# /etc/rsyslog.d/60-remote.conf  (every host)

template(name="RemoteFormat" type="string"
         string="%timereported% %HOSTNAME% %syslogtag%%msg%\n")

ruleset(name="sendToFiles") {
    action(
        type="omfwd"
        target="x.x.x.x"          # placeholder: the collector's address
        port="514"
        protocol="tcp"
        template="RemoteFormat"
        action.resumeRetryCount="100"  # keep retrying if the collector is down
        queue.type="linkedlist"        # queue in memory rather than dropping
        queue.size="10000"
        queue.saveonshutdown="on"      # persist the queue across a restart
    )
}

The queue settings matter more than they look. action.resumeRetryCount and the linked-list queue mean that if the collector is down for a while, the sending host holds its logs and delivers them when the collector returns, rather than silently dropping them. For a monitoring system, silently losing the events you are trying to monitor is the worst failure, so the forwarding is configured to tolerate an outage without data loss.

With the ruleset defined, each host calls it for the log streams worth forwarding. Rather than shipping every log line from every host (which would drown the collector in noise), each host forwards the categories that matter: high-severity events, authentication activity, boundary and firewall events, health snapshots, and a few service-specific streams. The next page is about which categories those are and why.

# Example: forward health snapshots and high-severity events
if $syslogtag startswith "health" then {
    call sendToFiles
    stop
}
if $syslogseverity <= 3 then {
    call sendToFiles
    stop
}
# ...further category rules follow the same shape

When a service logs to files, not syslog

The forwarding above assumes a host's programs log through the system logger, which most system services do. But plenty of applications write their own log files instead and never touch syslog: web servers, containerized apps, and many self-hosted services log to a path under their own directory. Those logs will not be collected by the rules above, because they never enter rsyslog in the first place.

The bridge is rsyslog's imfile module, which tails a file and injects each new line into rsyslog as though it had been logged normally, with a tag and facility you assign. Once a file-based log is tagged this way, it flows through the same forwarding rules as everything else. This is how a service that logs to files gets pulled into the pipeline without changing the application.

A per-service custom config on the host that produces the logs (kept in its own file so it is easy to find and change) does the tailing:

# /etc/rsyslog.d/61-custom.conf  (on the host running the service)
module(load="imfile")
 
# A service's auth log (e.g. failed logins, MFA events).
input(
  type="imfile"
  File="/opt/<service>/logs/identity/*.txt"   # the app's own log path
  Tag="myservice-identity"                     # tag it so rules can match it
  Severity="info"
  Facility="local6"
  PersistStateInterval="100"                   # remember read position periodically
  freshStartTail="on"                          # on first start, read new lines only,
                                               # not the entire existing file
)
 
# The same service's web access log (e.g. registrations, requests).
input(
  type="imfile"
  File="/opt/<service>/logs/nginx/access.log"
  Tag="myservice-nginx"
  Severity="info"
  Facility="local6"
  PersistStateInterval="100"
)

Two options are worth understanding:

  • PersistStateInterval makes rsyslog remember how far it has read into each file, so a restart resumes where it left off rather than re-reading (and re-forwarding) everything.
  • freshStartTail="on" means that the very first time rsyslog sees the file, it starts from the end rather than ingesting the entire backlog. Without it, a first start floods the collector with the whole history of the file. Use it on high-volume logs where the backlog is noise; omit it where you genuinely want the existing contents ingested once. With the file tagged, the collector's routing files it like any other tagged stream, and the forwarding rules send it on. The tag is the seam: everything downstream, the per-host log file on the collector, the analysis scripts, treats a tailed file-based log exactly like a native syslog stream, because by the time it leaves the host it is one.

A note on what is forwarded

Not everything is forwarded, and that is deliberate. Forwarding every log line would bury the signal in volume and strain the collector. The selection, which categories get sent, is a design decision, and the subject of the next page. The collection mechanism shown here is neutral plumbing; the judgment is in what you choose to run through it.

Adapt this for…

Any environment that needs central logging, which is any environment with more than one host. The tool is rsyslog and the transport is plain syslog over TCP, but the shape is universal; one collector with a listener and per-source separation, every host forwarding a selected set of streams with a queue that tolerates the collector being briefly down. Get the paths predictable and the forwarding loss-tolerant, and the analysis stage has a solid foundation to read from.