Skip to main content

Operating It

A WAF is not a set-and-forget install. It is a control you have to be able to see working, tune when it is wrong, and trust when it fires. This page is about running it after it is standing.

Logging

A control you cannot observe is a control you cannot trust. Open-appsec's logging is driven by the log trigger configured in the policy. Events, requests inspected, requests blocked, what rule or model score triggered a block, are logged locally.

Logs that live only on the WAF are lost if the WAF is the thing that fails, and they cannot be correlated with what the backends saw. When centralised with the rest of your host, application, and service logs, they become part of one picture. How I manage log forwarding, centralizing, parsing, and alerting, will be covered in a future deep dive.

The two questions the logs exist to answer:

  • What is being blocked? A blocked legitimate request is a person locked out. A blocked malicious request is the WAF doing its job and you want to confirm it is happening.
  • What is being flagged but not blocked? The protections in learning mode are logging what they would block. That stream is how you decide whether a protection is ready to promote to blocking.

Telling a real attack from a false positive

This is a core operational skill, and it is a judgment call the logs inform rather than make for you. When a request is blocked, the question is whether it was a genuine attack or a legitimate request that looked like one.

  • A false positive typically comes from a known source, correlates with a real user action, and involves an application's own legitimate but unusual behaviour, a large structured upload, an API call with an odd-looking payload, a automation client's request shape. The fix is a narrow, documented exception or promoting the relevant protection's tuning, not turning inspection off.
  • A real attack typically comes from an unfamiliar source, does not correlate with any legitimate user action, and matches attack shapes, injection strings, traversal sequences, probes for known-vulnerable paths. The right response is to confirm the block worked and, if the source is persistent, consider blocking it further upstream.

The reason the conservative minimum-confidence: critical setting from the inspection page matters here is that it biases toward fewer false positives, which keeps this triage manageable for a solo operator. The tradeoff is some lower-confidence suspicious traffic is logged rather than blocked. This is deliberate. A flood of false positives trains you to ignore the alerts, which is worse than a slightly more permissive block threshold.

Alerting

Raw logs are necessary but not sufficient. You also need to be told when something warrants attention rather than having to go looking. In my environment, an automation workflow reviews the forwarded logs and raises an alert on patterns worth a human's attention, for example, a spike in blocked requests or repeated hits from one source. The same log-and-alert approach covers the identity provider's authentication events, so the WAF's alerting is one instance of a general pattern: forward the logs, let an automation watch them, and surface only what matters.

The principle, consistent across this environment: monitor the control working, not just its output. A WAF that has silently stopped inspecting, because an nginx upgrade broke the attachment, because a service died, looks exactly like a WAF that is inspecting and finding nothing. The way you tell them apart is by watching for the absence of the logs you expect, not only the presence of alarming ones. A sudden silence from a component that normally logs steadily is itself a signal.

Testing that inspection actually works

Before trusting the WAF, and periodically after, confirm it is actually seeing and flagging malicious-looking requests. You do this by sending requests that mimic common attacks and checking that they show up in the logs (in detect-learn) or are refused (once promoted to prevent-learn). These are safe to run against your own environment.

Run them from an external host, not from an exempted source, or the exception will skip inspection and you will learn nothing.

SQL injection in a query parameter:

curl -k "https://app.example.org/?id=1' OR '1'='1"
curl -k "https://app.example.org/?id=1;DROP TABLE users--"

Cross-site scripting in a parameter:

curl -k "https://app.example.org/?q=<script>alert(1)</script>"

Path traversal, attempting to escape the web root:

curl -k "https://app.example.org/../../../../etc/passwd"
curl -k "https://app.example.org/?file=../../../../etc/passwd"

Command injection in a parameter:

curl -k "https://app.example.org/?host=127.0.0.1;cat%20/etc/passwd"

A non-standard HTTP method (this one is blocked outright by non-valid-http-methods, even early, since it is a boolean rather than a learned protection):

curl -k -X BADMETHOD "https://app.example.org/"

A known-scanner user agent, the kind of probe that hits every public host:

curl -k -A "sqlmap/1.0" "https://app.example.org/"

What to expect at each stage:

  • In detect-learn: these requests are not blocked. The application responds normally (or with its own 404), but each should appear in the open-appsec logs as something the WAF would have blocked. If they do not appear in the logs at all, inspection is not working, the module did not load, the agent is not running, or the request never reached the WAF.
  • In prevent-learn: the malicious requests should receive the 403-forbidden response instead of reaching the application. Seeing the 403 is the confirmation that blocking is live.

Running this small battery after install (to confirm detection works) and again after promoting a protection (to confirm blocking works) turns "I think the WAF is protecting us" into "I watched it catch these." It is also a good periodic check, since a WAF that has silently stopped inspecting will let all of these through with no log entry.

Adding exceptions when legitimate traffic trips inspection

Reviewing the detect-learn logs will surface legitimate requests the model flags such as a backup client's sync traffic, an app's large structured upload, an automation tool's unusual request shape. Before promoting the relevant protection to blocking, these need an exception, or promotion will start legitimate traffic.

An exception goes in the default policy's exceptions block. Keep the condition as narrow as the situation allows, a specific source IP is far safer than a broad range:

    exceptions:
      - name: allow-backup-sync
        condition:
          source-ip: "203.0.113.10"      # the specific trusted source
          url: "/sync"                   # scope appropriately, don't blanket exempt an IP
        action: allow
     # exempt a specific endpoint rather than a whole source
      - name: allow-large-upload-path
        condition:
          url: "/api/upload"
        action: allow
      # exempt a parameter that legitimately carries markup
      - name: allow-html-body-field
        condition:
          paramName: "post_body"
        action: skip

Apply it with open-appsec-ctl --apply-policy, then re-run the legitimate traffic and confirm it is no longer flagged. Every exception is a standing hole in the inspection, so each one should be as specific as possible, named clearly, documented as to why it exists, and revisited periodically to confirm it is still needed. An exception whose reason no one remembers is a liability sitting inside your security control.

Promoting protections to prevent-learn

Once a protection has a few weeks, or possibly several months depending on your ecosystems overall traffic, of detection behind it and you have added exceptions for the legitimate traffic it flags, promote it to blocking. Do this one protection at a time, so that if a promotion starts refusing legitimate requests, you know exactly which change caused it and can revert just that one. Flipping everything to prevent-learn at once and hoping is precisely the mistake the learning phase exists to prevent.

A suggested order, with reasoning:

  1. non-valid-http-methods is effectively already enforcing (it is a boolean, not a learned protection) and has essentially no false-positive risk, since legitimate clients do not send malformed methods. It is the safe first thing to have blocking, and confirms your custom-response and the block path work end to end.
  2. web-attacks next. It is the protection you most want enforcing, injection, XSS, traversal, the core of what a WAF is for, and with minimum-confidence: critical it is tuned conservatively to keep false-positive rate low. Promote the web-attacks override-mode to prevent-learn, and its sub-protections (csrf-protection, error-disclosure, open-redirect) as each proves reliable in the logs.
  3. snort-signatures once you have reviewed what it flags. Signature matching can be noisier against real application traffic, so it benefits from a longer look before blocking, but it is worth enforcing once tuned.
  4. openapi-schema-validation only where you actually have a defined API schema for the traffic. Without one it has little to enforce, and promoting it can reject legitimate requests that simply do not match an assumed shape. Promote deliberately, per the applications that warrant it.
  5. anti-bot last, and most cautiously. Bot detection has the highest false-positive potential of the set, since legitimate automation, monitoring, and sync clients can look bot-like. Spend the most detection time here and promote it only once you are confident it will not lock out your own tooling.

After each promotion, re-run the relevant curl tests from the section above to confirm blocking is live, and watch the logs for a few days for legitimate traffic newly caught. The whole sequence will likely span several months from install to a fully-enforcing policy.