October 4, 2026 · Yunus Emre Vurgun

How Do awk, sed, and grep Work? A Practical Guide

cli · tutorial · devops · regex

grep finds lines, sed edits them, and awk processes them as structured records: together the three tools handle most log analysis, config edits, and CSV chores without writing a script. Learn one representative one-liner for each and you will reach for them constantly; learn the dozen most-used flags and you will rarely need anything else. This guide covers exactly that working set, with examples you can run as-is.

Which tool for which job

The division of labor is simple once stated plainly. Use grep when you want to select lines, sed when you want to transform text line by line, and awk when lines have fields — columns, CSV values, log parts — that you want to compute on. Piping them together covers jobs none of them handles elegantly alone.

ToolJobCanonical one-liner
grepFind lines matching a patterngrep -rn "TODO" src/
sedSubstitute, delete, or print linessed -i 's/http:/https:/g' *.html
awkField-aware processing and reportsawk -F, '{sum += $3} END {print sum}' data.csv

If you are new to the terminal generally, the Unix file commands cheat sheet covers the navigation half of the workflow; the text processing commands reference lists these tools alongside their companions like cut, sort, and uniq.

grep: finding lines

grep prints the lines that match a pattern, and its flags control what "match" means and what gets printed. The options below cover the overwhelming majority of real usage; -r for recursive search and -i for case-insensitivity lead by a wide margin.

FlagDoesExample
-rSearch directories recursivelygrep -r "password" config/
-iIgnore casegrep -i "error" app.log
-vInvert: show non-matching linesgrep -v "^#" settings.conf
-nShow line numbersgrep -n "return" main.py
-cCount matches instead of printinggrep -c "404" access.log
-EExtended regex (+ ? | ( ))grep -E "4[0-9]{2}" access.log
-oPrint only the matched partgrep -o "[0-9.]*ms" bench.log
-A/-B/-CShow lines after/before/around a matchgrep -C3 "Traceback" app.log

Two habits pay off immediately. First, quote your pattern — grep -r "*.tmp" . without quotes lets the shell expand the glob before grep sees it, producing nonsense. Second, pair -v with a second grep to carve signal from noise: grep "ERROR" app.log | grep -v "timeout" shows errors worth reading right now. For huge trees, grep -r --include="*.py" restricts extensions, though at that point many people switch to ripgrep (rg), which respects .gitignore and is dramatically faster.

sed: editing streams

sed reads input line by line, applies commands, and prints the result — a non-interactive editor ideal for bulk edits across files. Substitution is the command you will use most, with deletion and selective printing close behind. Full syntax lives in the sed reference; the essentials fit in one table.

CommandDoesExample
s/old/new/Replace first match per linesed 's/cat/dog/' pets.txt
s/old/new/gReplace all matches per linesed 's/;/,/g' data.txt
-iEdit files in placesed -i 's/v1/v2/g' *.md
dDelete matching linessed '/^$/d' notes.txt
-n 'Np'Print only line Nsed -n '10p' data.csv
-n '/re/p'Print only matching linessed -n '/ERROR/p' app.log
N,M commandRestrict to a line rangesed -n '5,15p' report.txt

Three warnings from experience. First, always test without -i before editing in place — sed prints to stdout by default, so a dry run costs nothing and a bad -i across hundreds of files costs plenty. Second, GNU and BSD (macOS) sed disagree on -i: GNU accepts a bare -i while BSD requires -i '', so scripts meant for both need the portable form. Third, the / delimiter is only a convention — when replacing paths, use another delimiter to dodge escaping: sed 's|/old/path|/new/path|g'.

awk: columns and reports

awk treats each line as a record split into fields — $1, $2, up to $NF (the last field) — and runs your program against every record. The pattern-action shape pattern { action } reads almost like English: match these lines, do this with them. The awk reference documents the full language; these five idioms handle most daily needs.

# Print columns 1 and 3 of a CSV file
awk -F, '{print $1, $3}' data.csv

# Sum column 3 and print the total at the end
awk -F, '{sum += $3} END {print sum}' data.csv

# Show lines where column 2 exceeds 100
awk '$2 > 100 {print $0}' measurements.txt

# Count requests per status code in an access log
awk '{count[$9]++} END {for (c in count) print c, count[c]}' access.log

# Reformat: swap first two columns, comma-separated output
awk -F, 'BEGIN {OFS=","} {print $2, $1, $3}' data.csv

The BEGIN block runs before input (set separators, print headers) and END runs after (print totals, summaries) — that trio of pattern, BEGIN, and END is the whole mental model. Note -F, sets the input separator; the default is any whitespace, which conveniently parses most logs. Associative arrays like count[$9] appear with zero setup and make awk a genuine one-line reporting engine.

Putting them together: three real pipelines

The tools shine in combination, each doing the job it does best. These pipelines are deliberately ordinary — the kind of thing that comes up weekly.

# Top 10 slowest API endpoints from a log (grep filters, awk reports)
grep "GET /api" access.log | awk '{print $7, $NF}' | sort | uniq -c | sort -rn | head

# Rename a function across a codebase (grep finds, sed rewrites)
grep -rl "old_name" src/ | xargs sed -i 's/old_name/new_name/g'

# Audit: uncommented settings and their values
grep -v "^#" app.conf | grep "=" | awk -F= '{print $1}' | sort

Read pipelines right to left when debugging: check what the last stage receives before blaming the last stage itself. And remember sort | uniq -c | sort -rn as a fixed idiom — frequency tables in three commands, endlessly reusable.

FAQ: should I learn awk or just use Python?

Should I learn awk or just use Python? Both, at different scales. For one-liners and quick log questions, awk wins: no file to create, no imports, instant answers. Once logic needs conditionals, error handling, JSON parsing, or reuse, Python wins — python3 -c bridges the gap for medium jobs. The practical rule: if the awk program needs more than one END block's worth of logic, switch languages.

Is grep obsolete given ripgrep? Not obsolete, but rg is worth installing: it is faster, respects .gitignore by default, and searches hidden things less often by accident. The flags differ slightly (rg -i, rg -v, rg -l), but the muscle memory transfers. grep remains the universal default on every Unix box you will ever ssh into.

What about perl one-liners? Perl's -pe and -ne flags make it a capable sed-and-awk replacement with full regex power, and some veterans prefer it. It is worth recognizing in others' scripts; it is rarely worth choosing for new work over Python.