Analysing Nginx Access Logs with awk

An access log is a plain text file with one request per line, which makes awk a quick way to answer questions like which pages are failing, who is hammering the server, or how much bandwidth a day used. The only difficulty is knowing which field is which.

Which field is which

In the default combined format awk splits each line on spaces, giving a fixed set of fields: $1 is the client IP, $4 starts the timestamp with a leading bracket, $6 is the request method with a leading quote, $7 is the path, $9 is the status code and $10 is the response size in bytes. The quoted request line contains spaces, which is why the status is field 9 and not field 7.

203.0.113.4 - - [04/Oct/2026:10:00:01 +0000] "GET /a HTTP/1.1" 200 1234 "-" "curl/8.4"
$1                $4                     $6   $7 $8        $9  $10

Filtering

A pattern before the braces selects lines, and conditions combine with && and ||. Comparing the status to a number matches an exact code, while a regular expression on the first digit matches a whole class. The method has a leading quote, so compare it with substr($6, 2).

awk '$9 ~ /^5/ { print }' access.log
awk 'substr($6, 2) == "POST" && $9 == 401' access.log

Ranking with a counter

An associative array counts occurrences by key, and an END block prints the totals once the file is read. Pipe the result through sort -rn and head to keep the top few. Cut the query string off a path first, or every different query counts as a separate page.

awk '{ c[$1]++ } END { for (k in c) print c[k], k }' access.log | sort -rn | head -n 10

Limits of this approach

This assumes the default combined format, so a custom log_format or JSON logs need different fields or a tool such as jq. A user agent or referer containing spaces does not disturb the fields used here, because they all come before it. Rotated logs are compressed, which is why reading all of them needs zcat -f.

Open the Access log filter generator

Frequently asked questions

Why is the status code field 9 and not 7?

The request in quotes holds three space-separated words: method, path and protocol. Awk counts each as its own field, so the status lands after them, at field 9.

How do I count requests per hour?

Take the hour from the timestamp field with substr($4, 14, 2) and count it in an array, as in the top-IP example. The first character of $4 is a bracket, so the hour starts at position 14.

Is awk fast enough for a large log?

Yes for most servers: it reads a line at a time and uses little memory, handling millions of lines in seconds. For very large or continuously growing logs, a log pipeline or GoAccess gives live dashboards.

Guides