Docker Deep Dive · Prerequisites · P3 of 7

Streams, Pipes & Text Tools

Every process is born holding three wires. Learn to reroute them, weld processes together with pipes, and interrogate a 500-line web log — which is exactly the machinery behind docker logs.

⏱ ~25 min hands-on · repo chapter p3-streams-pipes-and-text · cheatsheet: Linux & Shell

prereq P3 / 7

Before you start

cd docker-deep-dive
git pull
cd p3-streams-pipes-and-text
docker run -it --name pipebox -v "$(pwd)":/labs ubuntu:24.04 bash
docker run -it --name pipebox -v "${PWD}:/labs" ubuntu:24.04 bash
cd /labs

The -v flag shares this chapter folder into the container at /labs (lesson 5 explains bind mounts properly — today it just gets the dataset inside). If you're doing the prerequisites first: that one command is a disposable Linux machine; exit leaves it, nothing on your computer can break.

The dataset, access.log, is one busy hour of traffic to Harbor Log — this course's app. Two incidents are buried in it. Your pipelines will find both.

1.Three wires soldered to every process

When Linux starts a process, it hands it three open streams, known by number: 0 — stdin (what it reads), 1 — stdout (its results), 2 — stderr (its complaints). In a terminal all three point at you: 0 at your keyboard, 1 and 2 at your screen.¹ Two separate output wires is the trick — results stay clean even when errors happen.

keyboard (your typing) a process e.g. ls, grep, python screen results screen complaints 0 · stdin 1 · stdout 2 · stderr
File descriptors 0, 1, 2. Every process gets these three at birth; redirection and pipes work by re-pointing them before the process starts. The numbers are the real names the kernel uses — 2> literally means "re-point wire 2".

Prove the two output wires are separate. Ask ls about one file that exists and one that doesn't:

ls /etc/hostname /nope
ls: cannot access '/nope': No such file or directory
/etc/hostname

Now capture the output into a file with > — and watch what doesn't get captured:

ls /etc/hostname /nope > out.txt
cat out.txt
ls: cannot access '/nope': No such file or directory   ← still on screen!
/etc/hostname                                          ← this is cat showing out.txt

> re-pointed wire 1 only. The complaint rode wire 2, which still aimed at your screen. This is by design: a script can save clean results while failures stay loud.²

2.Rerouting: redirection

syntaxre-pointsto
cmd > fstdout (1)file f — truncating it first
cmd >> fstdout (1)file f — appending
cmd 2> fstderr (2)file f
cmd > f 2>&1bothone file (2 follows 1 — order matters)
cmd < fstdin (0)read from file instead of keyboard
cmd 2> /dev/nullstderr (2)the black hole — a real file (P2!) that discards everything

> truncates before the command even runs

The shell empties the target file first, then starts the command. So sort data.txt > data.txt destroys your data — the file is empty before sort reads a byte. Redirect to a new file, or use >> when adding.²

Drills, on the dataset:

grep " 500 " access.log > errors.txt    # server errors → their own file
wc -l errors.txt                        # how many lines landed there?
grep " 404 " access.log >> errors.txt   # append the not-founds
wc -l errors.txt
ls /labs /nope 2> /dev/null             # complaints, discarded
ls /labs /nope > everything.txt 2>&1    # both wires → one file
cat everything.txt
rm out.txt errors.txt everything.txt    # leave the lab tidy
42 errors.txt
72 errors.txt

42 server errors, plus 30 not-founds = 72. Hold those numbers — you're about to find out where they came from.

3.Pipes: weld processes together

A pipe | connects one process's wire 1 to the next process's wire 0, in memory, live — no temp files. It's the heart of the Unix idea, as Doug McIlroy (the pipe's inventor) put it: "Make each program do one thing well… expect the output of every program to become the input to another."³ Small tools, composed on demand, answer questions no single tool anticipated.

Investigation 1: Harbor Log is erroring — which endpoint is failing? Build the answer one stage at a time, checking each stage's output before adding the next. Stage 1 — keep only the error lines (" 500 " with spaces, so byte counts can't false-match):

grep " 500 " access.log | head -2      # peek at what matched
grep " 500 " access.log | wc -l        # how many matched
192.168.65.1 - - [17/Jul/2026:09:10:55 +0000] "GET /api/messages HTTP/1.1" 500 1102
198.51.100.42 - - [17/Jul/2026:09:15:12 +0000] "GET / HTTP/1.1" 500 1102
42

Stage 2 — keep only the URL. Split each line on spaces (-d' ') and take field 7 (-f7):

grep " 500 " access.log | cut -d' ' -f7 | head -4
/api/messages
/
/api/messages
/api/upload

Stage 3 — group and rank. uniq -c collapses repeated lines and counts them, but it only sees adjacent repeats — so sort must run first to bring duplicates together. Then sort the counts, biggest first:⁴

grep " 500 " access.log | cut -d' ' -f7 | sort | uniq -c | sort -rn | head -3
     37 /api/upload
      3 /api/messages
      2 /

Verdict in one line: /api/upload threw 37 of the 42 errors — that's where you'd start debugging. sort | uniq -c | sort -rn is the "group-and-rank" idiom; it will answer questions for you for the rest of your career.

access.log 500 lines grep " 500 " wire 1 out of grep — only matching lines survive: 192.168.65.1 … "GET /api/messages HTTP/1.1" 500 1102 198.51.100.42 … "GET / HTTP/1.1" 500 1102 10.0.4.23 … "POST /api/upload HTTP/1.1" 500 1523 …39 more — 42 of 500 lines remain cut -d' ' -f7 each line loses everything but field 7 — the path: /api/messages / /api/upload …still 42 lines, but now one path per line sort alphabetical order — duplicates become ADJACENT: / / /api/messages /api/messages /api/messages /api/upload /api/upload /api/upload /api/upload … this is why sort must come before uniq uniq -c count runs each run of identical lines collapses to count + line: 2 / 3 /api/messages 37 /api/upload 42 lines became 3 groups sort -rn | head -3 ranked, biggest first — the answer: 37 /api/upload ← debug here 3 /api/messages 2 / the shell started all 5 processes AT ONCE — the kernel streams bytes between them, no temp files each | is wire 1 of the left process soldered to wire 0 of the right
  1. grep starts, reading the log on wire 0 and emitting only lines containing " 500 " on wire 1 — 42 of 500 survive.
  2. cut receives those 42 lines and keeps only field 7 of each: the URL path.
  3. sort buffers, then emits the paths in order — identical paths are now adjacent.
  4. uniq -c collapses each adjacent run into "count line": three groups.
  5. sort -rn ranks numerically, biggest first; head -3 keeps the podium. Verdict: /api/upload, 37 errors.
  6. And the secret: all five processes ran concurrently — a pipe is plumbing, not a sequence of saves.
The pipeline as an assembly line. Press ▶ to watch the data shrink at each station: 500 log lines → 42 matches → 42 paths → 3 groups → 1 verdict. Each | connects wire 1 to wire 0; the kernel does the plumbing.¹

Your turn — investigation 2: who's scanning us?

The log also holds a probe: someone requesting /wp-login.php, /.env, /admin… on a Flask app that has none of them. Same pipeline shape, different filter and field: 404s, field 1 (the client IP). Build it yourself, then grade all three answers:

grep " 404 " access.log | wc -l
# …your pipeline here: which IP owns most of those 404s?
bash check.sh <count-of-500s> <top-failing-path> <scanner-ip>
OK  count of 500s   : 42
OK  top failing path: /api/upload
OK  scanner IP      : 203.0.113.66
PASS — pipeline skills confirmed.

Useful grep variations while you work:⁵

flagdoestry
-ccount matches instead of printinggrep -c " 404 " access.log
-vinvert: keep NON-matching linesgrep " 404 " access.log | grep -v 203.0.113.66
-nshow line numbersgrep -n "upload" access.log | head -3
-iignore casegrep -i "get" access.log | wc -l
-rrecurse through a directorygrep -r "PATH" /etc/profile.d 2>/dev/null

Watch a file grow: tail -f

Real logs don't hold still. tail -f prints the end of a file and then follows it, streaming new lines as they land. Two terminals — in the first (inside pipebox):

tail -f access.log

In a second terminal on your machine, open another shell into the same container (lesson 1's docker exec) and append a line:

docker exec -it pipebox bash
echo '10.0.4.99 - - [17/Jul/2026:10:00:00 +0000] "GET / HTTP/1.1" 200 1234' >> /labs/access.log
192.168.65.1 - - [17/Jul/2026:09:59:58 +0000] "GET / HTTP/1.1" 200 1981
10.0.4.99 - - [17/Jul/2026:10:00:00 +0000] "GET / HTTP/1.1" 200 1234   ← appeared live

Ctrl-C stops the follow; tidy up with sed -i '/10.0.4.99/d' access.log. You already know this feeling: docker logs -f is tail -f on a container's output — same idea, same Ctrl-C.

4.find — the other search

grep looks inside files; find looks for files — walking a directory tree and filtering by name, type, age, or size:⁶

find /etc -name '*.conf' -type f | head -5   # config files under /etc (P2's tour!)
find /labs -type f -mtime -1                 # files modified in the last day
find / -size +5M -type f 2>/dev/null | head  # big files, complaints discarded
/etc/sysctl.conf
/etc/libaudit.conf
/etc/host.conf
/etc/sysctl.d/10-console-messages.conf
/etc/sysctl.d/10-zeropage.conf

Two ways to feed one command's output into another command's arguments (different from piping into its stdin!) — you met the first in lesson 1's cleanup:

docker rm $(docker ps -aq)        # $( ) = run inner command, paste its output HERE
docker ps -q | xargs docker stop  # xargs = turn stdin lines into arguments

5.Why Docker cares about all of this

Here's the payoff. A container's main process gets the same three wires — but Docker captures wires 1 and 2 instead of pointing them at a screen. docker logs simply replays the recording. Watch both wires get caught:⁷

docker run --name logdemo ubuntu:24.04 bash -c 'echo "harbor log ready"; echo "disk almost full" >&2'
docker logs logdemo
docker rm logdemo
disk almost full     ← rode wire 2
harbor log ready     ← rode wire 1 — both were captured

That's why 12-factor apps log to stdout and never to log files: the process just writes to the wire it was born with, and the platform — Docker today, Kubernetes in your next course (kubectl logs) — owns collection, routing, and storage.⁸ Harbor Log has done this since lesson 3 without you noticing.

And your new pipeline skills compose straight onto it:

docker logs web 2>&1 | grep ERROR | wc -l   # logs replays on both wires → merge, filter, count

You can now

6.Check yourself

What does cmd > results.txt do to an existing results.txt before cmd runs?

  1. empties it, then writes
  2. keeps it, then appends
  3. asks before overwriting it
  4. fails, refusing to overwrite

You run app > out.txt and an error message still appears on your screen. Why?

  1. buffering delayed the output
  2. the redirect silently failed
  3. errors travel on stderr
  4. the file was read-only

Why must sort run before uniq -c?

  1. sort removes all duplicate lines
  2. uniq only collapses adjacent duplicates
  3. uniq requires numerically sorted input
  4. pipes shuffle the line order

Write the pipeline that answers: "which endpoint threw the most 500s in access.log?"

grep " 500 " access.log | cut -d' ' -f7 | sort | uniq -c | sort -rn | head -3 — filter the error lines, keep the path field, sort so duplicates touch, count each group, rank numerically, keep the top. (head -1 also fine.)

Your container wrote nothing to disk — so what is docker logs showing you, exactly?

The captured stdout and stderr (wires 1 and 2) of the container's main process. Docker records both streams instead of pointing them at a terminal; logs replays the recording, -f follows it live — which is why 12-factor apps just write to stdout and let the platform collect.

7.Go deeper

Primary source: The Linux Command Line (free book), chapters 6–7 — "Redirection" and "Seeing the World as the Shell Sees It" — cover today's ground gently and thoroughly. The precise rules live in the Bash manual's Redirections section.

Where this lands in the course: lesson 1's cleanup one-liner and docker logs, lesson 3 & 8's RUN pipelines (apt-get update && apt-get install…, curl … | tar xz), and lesson 12's log-based debugging all assume today's plumbing. Next: P4 — Processes, Environments & Signals, where the things holding these wires — processes — get names, parents, environments, and a polite way to die.

Stuck? Curious?

Bring questions to class, or open an issue on the course repo — include the command you ran and the output you got. The quizzes above are for self-checking: commit to an answer before revealing it, and re-try anything you missed tomorrow.

Sources

  1. Shotts — The Linux Command Line, ch. 6–7 (streams, redirection, pipelines)
  2. GNU Bash Reference Manual — Redirections (truncation, 2>&1 ordering)
  3. McIlroy, Pinson & Tague — BSTJ 57(6), 1978 foreword ("do one thing well… output becomes input")
  4. GNU Coreutils manual (cut, sort, uniq, wc, tail)
  5. GNU Grep manual (flags)
  6. GNU Findutils manual (find tests and actions)
  7. Docker Docs — docker logs (captures STDOUT/STDERR)
  8. The Twelve-Factor App — XI. Logs (logs as event streams to stdout)