Puzzleshot #020 - REDACTED (with sed)
We’ve used awk, DuckDB, Python, and jq throughout this series — but there’s a classic unix tool we haven’t touched yet: sed.
This puzzle’s scenario: you need to share a log file externally — maybe for a vendor ticket, a bug report, or a support request — and before you send it out you need to redact any IP addresses it contains.
Sample log data:
2026-04-09 10:15:32 GET /api/orders 200 client=192.168.1.45
2026-04-09 10:15:33 GET /api/dealers 200 client=10.0.0.12
2026-04-09 10:15:34 POST /api/orders 500 client=203.0.113.7
2026-04-09 10:15:35 GET /api/orders 200 client=192.168.1.46
The challenge:
Using sed, redact every IPv4 address in the file, replacing each one with the placeholder REDACTED_IP.
Bonus: instead of full redaction, try preserving just the first octet (ie, 192.xxx.xxx.xxx) for partial anonymization.
Things to consider:
- What regex pattern correctly matches an IPv4 address?
- Could this pattern produce false positives — matching something that looks like an IP but isn’t?
- What’s the difference between basic sed and sed -E (extended regex), and which do you need here?
- When would you reach for sed over awk for a task like this?
Reveal solution
This puzzle we used sed to redact IP addresses from a log file before sharing it externally.
Full redaction:
sed -E 's/[0-9]{1,3}\.[0-9]{1,3}\.[0-9]{1,3}\.[0-9]{1,3}/REDACTED_IP/g' access.log
What’s happening:
[0-9]{1,3} matches one to three digits — a single octet. Repeated four times with literal dots (\.) in between, this matches the full dotted-decimal IPv4 pattern. The g flag ensures every match on each line gets replaced, not just the first.
-E enables extended regex syntax, which lets you write {1,3} directly. Without -E (basic sed) you’d need to escape the braces as \{1,3\} instead — both work, but -E is more readable for anyone used to regex in other languages.
Bonus — preserving the first octet:
sed -E 's/([0-9]{1,3})\.[0-9]{1,3}\.[0-9]{1,3}\.[0-9]{1,3}/\1.xxx.xxx.xxx/g' access.log
Wrapping the first octet in parentheses captures it as a group. \1 in the replacement refers back to that captured value (a back-reference, common & useful in regex), so 192.168.1.45 becomes 192.xxx.xxx.xxx — enough anonymization to hide the specific host while preserving which broad network range it came from.
The false positive gotcha:
This regex isn’t specific to valid IPs — it matches any four dot-separated groups of one to three digits. That means something like a version number gets caught in the same net:
echo "Deployed app version 12.4.1.0 to server client=192.168.1.45" | sed -E 's/[0-9]{1,3}\.[0-9]{1,3}\.[0-9]{1,3}\.[0-9]{1,3}/REDACTED_IP/g'
Both the version number and the actual IP get redacted. Whether that matters depends on your data — if your log files never contain version strings in that format you’re fine, but it’s worth knowing the regex is pattern-matching, not IP-validating. A stricter regex bounding each octet to 0-255 is possible but considerably messier, and in practice most people accept the simpler pattern and verify their data doesn’t collide with it.
Why sed over awk here:
This is a straightforward whole-line text substitution with no need to reference specific fields or columns. sed is built exactly for this — pattern in, pattern out, one line of logic. awk would work too but requires wrapping the same regex in a print/gsub structure for no real benefit when you’re not doing anything field-aware.