Puzzleshot #023 - Tailed streams with awk
So far every puzzle has started with a finished file. Real systems don’t wait for you — the log is still being written while you’re trying to read it.
This scenario: a service is writing web logs to access.log right now, and you want a running tally of requests per minute (plus how many were server errors) printed as each minute completes. No Kafka, no stream processor, no database — just the unix toolbox.
Sample log lines (date, time, method, path, status):
2026-10-08 10:00:01 GET /api/orders 200
2026-10-08 10:00:03 GET /api/dealers 200
2026-10-08 10:00:04 POST /api/orders 500
2026-10-08 10:01:02 GET /api/orders 200
The challenge:
Using tail and awk, print one line per minute — the minute, the total number of requests, and how many were 5xx errors — as each minute completes. Something like:
10:00 30 1
10:01 31 3
Things to consider:
- What does tail -f actually follow, and what happens when the log gets rotated?
- Your script works at the terminal. Does it still work when you pipe its output into a file or another command?
- When does a minute’s line actually get printed? What happens to the last minute if traffic stops?
Reveal solution
This puzzle we turned a log that’s still being written into per-minute counts, using nothing but tail and awk.
A solution:
tail -n +1 -F access.log | awk '
{
m = substr($2, 1, 5)
if (m != cur) {
if (cur != "") { print cur, total, errs; fflush() }
cur = m; total = 0; errs = 0
}
total++
if ($5 >= 500) errs++
}'
What’s happening:
tail -n +1 -F starts at the first line of the file and keeps following it. The capital F matters: it follows the file by name, so when the log is rotated (renamed, with a fresh file started in its place) tail notices and switches over. Plain -f follows the open file handle instead, so after a rotation it keeps watching the old, renamed file and goes quiet. In testing, -f never showed a line from the new file, while -F reported that the file had reappeared and carried on.
awk takes the HH:MM from the time field and counts requests as they arrive. When the minute changes, the previous minute must be complete, so it prints that row and starts a fresh count. $5 >= 500 counts the server errors.
fflush() is the piece that bites people. When awk’s output goes to a terminal you see it as it’s printed. When it goes to a pipe or a file, awk buffers it, and nothing downstream sees a thing until the buffer fills or awk exits. In testing, a version without fflush() delivered zero bytes while it ran, and the whole output only appeared once I stopped tail.
There’s a second buffering layer on some systems. mawk, the default awk on some Debian and Ubuntu installs, also buffers its input from a pipe, so even with fflush() it stayed silent in testing. mawk -W interactive fixed it, and so did gawk. It’s worth checking which awk you actually have (awk –version, or awk -W version).
The last minute:
A minute’s row only prints when a line from a later minute arrives, because that’s the only way awk knows the minute is over. If traffic stops, the final minute never prints. In my test run, 120 lines were logged but the printed rows covered only 91 of them. awk has no timer, it only wakes up when a line shows up. Streaming engines face the same question — when is it safe to close a window? — and answer it with watermarks and timeouts.
When this is enough, and when it isn’t:
For one box and one metric, a tail and an awk script is hard to beat. It starts to hurt with several servers (you’d have to merge the streams), when you need to survive a restart (tail keeps no offsets, so awk’s counts are lost if it dies), or when events can arrive out of order.