October 4, 2026 · Yunus Emre Vurgun
How Do awk, sed, and grep Work? A Practical Guide
grep finds lines, sed edits them, and awk processes them as structured records: together the three tools handle most log analysis, config edits, and CSV chores without writing a script. Learn one representative one-liner for each and you will reach for them constantly; learn the dozen most-used flags and you will rarely need anything else. This guide covers exactly that working set, with examples you can run as-is.
Which tool for which job
The division of labor is simple once stated plainly. Use grep when you want to select lines, sed when you want to transform text line by line, and awk when lines have fields — columns, CSV values, log parts — that you want to compute on. Piping them together covers jobs none of them handles elegantly alone.
| Tool | Job | Canonical one-liner |
|---|---|---|
grep | Find lines matching a pattern | grep -rn "TODO" src/ |
sed | Substitute, delete, or print lines | sed -i 's/http:/https:/g' *.html |
awk | Field-aware processing and reports | awk -F, '{sum += $3} END {print sum}' data.csv |
If you are new to the terminal generally, the Unix file commands cheat sheet covers the navigation half of the workflow; the text processing commands reference lists these tools alongside their companions like cut, sort, and uniq.
grep: finding lines
grep prints the lines that match a pattern, and its flags control what "match" means and what gets printed. The options below cover the overwhelming majority of real usage; -r for recursive search and -i for case-insensitivity lead by a wide margin.
| Flag | Does | Example |
|---|---|---|
-r | Search directories recursively | grep -r "password" config/ |
-i | Ignore case | grep -i "error" app.log |
-v | Invert: show non-matching lines | grep -v "^#" settings.conf |
-n | Show line numbers | grep -n "return" main.py |
-c | Count matches instead of printing | grep -c "404" access.log |
-E | Extended regex (+ ? | ( )) | grep -E "4[0-9]{2}" access.log |
-o | Print only the matched part | grep -o "[0-9.]*ms" bench.log |
-A/-B/-C | Show lines after/before/around a match | grep -C3 "Traceback" app.log |
Two habits pay off immediately. First, quote your pattern — grep -r "*.tmp" . without quotes lets the shell expand the glob before grep sees it, producing nonsense. Second, pair -v with a second grep to carve signal from noise: grep "ERROR" app.log | grep -v "timeout" shows errors worth reading right now. For huge trees, grep -r --include="*.py" restricts extensions, though at that point many people switch to ripgrep (rg), which respects .gitignore and is dramatically faster.
sed: editing streams
sed reads input line by line, applies commands, and prints the result — a non-interactive editor ideal for bulk edits across files. Substitution is the command you will use most, with deletion and selective printing close behind. Full syntax lives in the sed reference; the essentials fit in one table.
| Command | Does | Example |
|---|---|---|
s/old/new/ | Replace first match per line | sed 's/cat/dog/' pets.txt |
s/old/new/g | Replace all matches per line | sed 's/;/,/g' data.txt |
-i | Edit files in place | sed -i 's/v1/v2/g' *.md |
d | Delete matching lines | sed '/^$/d' notes.txt |
-n 'Np' | Print only line N | sed -n '10p' data.csv |
-n '/re/p' | Print only matching lines | sed -n '/ERROR/p' app.log |
N,M command | Restrict to a line range | sed -n '5,15p' report.txt |
Three warnings from experience. First, always test without -i before editing in place — sed prints to stdout by default, so a dry run costs nothing and a bad -i across hundreds of files costs plenty. Second, GNU and BSD (macOS) sed disagree on -i: GNU accepts a bare -i while BSD requires -i '', so scripts meant for both need the portable form. Third, the / delimiter is only a convention — when replacing paths, use another delimiter to dodge escaping: sed 's|/old/path|/new/path|g'.
awk: columns and reports
awk treats each line as a record split into fields — $1, $2, up to $NF (the last field) — and runs your program against every record. The pattern-action shape pattern { action } reads almost like English: match these lines, do this with them. The awk reference documents the full language; these five idioms handle most daily needs.
# Print columns 1 and 3 of a CSV file
awk -F, '{print $1, $3}' data.csv
# Sum column 3 and print the total at the end
awk -F, '{sum += $3} END {print sum}' data.csv
# Show lines where column 2 exceeds 100
awk '$2 > 100 {print $0}' measurements.txt
# Count requests per status code in an access log
awk '{count[$9]++} END {for (c in count) print c, count[c]}' access.log
# Reformat: swap first two columns, comma-separated output
awk -F, 'BEGIN {OFS=","} {print $2, $1, $3}' data.csvThe BEGIN block runs before input (set separators, print headers) and END runs after (print totals, summaries) — that trio of pattern, BEGIN, and END is the whole mental model. Note -F, sets the input separator; the default is any whitespace, which conveniently parses most logs. Associative arrays like count[$9] appear with zero setup and make awk a genuine one-line reporting engine.
Putting them together: three real pipelines
The tools shine in combination, each doing the job it does best. These pipelines are deliberately ordinary — the kind of thing that comes up weekly.
# Top 10 slowest API endpoints from a log (grep filters, awk reports)
grep "GET /api" access.log | awk '{print $7, $NF}' | sort | uniq -c | sort -rn | head
# Rename a function across a codebase (grep finds, sed rewrites)
grep -rl "old_name" src/ | xargs sed -i 's/old_name/new_name/g'
# Audit: uncommented settings and their values
grep -v "^#" app.conf | grep "=" | awk -F= '{print $1}' | sortRead pipelines right to left when debugging: check what the last stage receives before blaming the last stage itself. And remember sort | uniq -c | sort -rn as a fixed idiom — frequency tables in three commands, endlessly reusable.
FAQ: should I learn awk or just use Python?
Should I learn awk or just use Python? Both, at different scales. For one-liners and quick log questions, awk wins: no file to create, no imports, instant answers. Once logic needs conditionals, error handling, JSON parsing, or reuse, Python wins — python3 -c bridges the gap for medium jobs. The practical rule: if the awk program needs more than one END block's worth of logic, switch languages.
Is grep obsolete given ripgrep? Not obsolete, but rg is worth installing: it is faster, respects .gitignore by default, and searches hidden things less often by accident. The flags differ slightly (rg -i, rg -v, rg -l), but the muscle memory transfers. grep remains the universal default on every Unix box you will ever ssh into.
What about perl one-liners? Perl's -pe and -ne flags make it a capable sed-and-awk replacement with full regex power, and some veterans prefer it. It is worth recognizing in others' scripts; it is rarely worth choosing for new work over Python.