12  Pipes & Redirection

So far every command has worked on one file at a time. Real data work chains small tools together: decompress a log, skip its header, pull out one column, count values, and keep the top 10 — without ever opening the 5,000-row file in an editor or loading it into R first.

12.1 Why pipes matter for R users

Before calling readr::read_csv() on a large download, you want to know: how many rows? Which columns? What are the most frequent values? The shell answers in milliseconds by streaming data through a pipeline. Nothing is stored in between — each command reads, transforms, and passes onward.

12.2 The | operator

The pipe | connects the standard output of one command to the standard input of the next:

gzip -dc cran_log_sample.csv.gz | head -n 3
"date","time","size","r_version","r_arch","r_os","package","version","country","ip_id"
"2019-09-15","02:28:38",47952,NA,NA,NA,"aws.s3","0.3.12","US",1
"2019-09-15","02:28:37",5899,NA,NA,NA,"aws.ec2metadata","0.2.0","US",1

gzip -dc decompresses to standard output (-d) and keeps the original file (-c), so the fixture stays compressed. We use gzip -dc instead of zcat because zcat fails on macOS (it expects .Z files there).

On macOS, several coreutils differ from Linux GNU versions (e.g. date, sed -i, head -n -5). The commands in this chapter behave identically on both.

12.3 The canonical pipeline

Our question: which packages were downloaded most in the CRAN log sample? Build the answer one stage at a time:

gzip -dc cran_log_sample.csv.gz | tail -n +2 | cut -d',' -f7 | tr -d '"' | sort | uniq -c | sort -nr | head -n 10
    124 rsconnect
    122 magrittr
    115 aws.s3
    108 aws.ec2metadata
     93 ggplot2
     47 stringr
     47 rlang
     45 dplyr
     44 testthat
     44 ellipsis

Stage by stage:

Stage Purpose
gzip -dc cran_log_sample.csv.gz Stream the uncompressed CSV (5,001 lines: header + 5,000 rows)
tail -n +2 Skip line 1, the header ("date","time",…), so "package" is not counted as a download
cut -d',' -f7 Extract field 7, the package column (field 3 is size — verify with the header)
tr -d '"' Strip quotes so "aws.s3" and aws.s3 group together
sort Group identical names (required before uniq -c)
uniq -c Count each group
sort -nr Sort numerically (-n), descending (-r)
head -n 10 Keep the top 10

The result: rsconnect (124 downloads) tops the sample, followed by magrittr (122) and aws.s3 (115).

cut -d',' splits on every comma naively — it breaks on CSVs with quoted commas inside fields. For such files use a real CSV parser (readr::read_csv() in R, or xsv/csvkit on the shell).

12.4 Redirection and chaining recap

Chapter 4 covered > (overwrite) and >> (append). Two more operators complete the set:

sort < pkg_names.txt > sorted_pkgs.txt

< feeds a file into a command’s standard input. And && runs the next command only if the previous one succeeded:

mkdir -p pipe_demo && cp release_names.txt pipe_demo/ && ls pipe_demo
release_names.txt
rm -rf pipe_demo

12.5 | versus |> and %>%

The shell pipe and the R pipe rhyme but differ:

Shell | R \|> / %>%
Passes raw text stream R objects
Evaluation streaming (constant memory) eager (each step completes first)
Errors exit codes (echo $?), easy to ignore silently conditions that stop unless handled

In R, the equivalent of the counting stage is table() + sort():

logs <- readr::read_csv("cran_log_sample.csv.gz", show_col_types = FALSE)
sort(table(logs$package), decreasing = TRUE)[1:10]

Use the shell to inspect and slim down data before it ever enters R; use R once the data fits comfortably in memory.

Command R
cmd1 \| cmd2 cmd2(cmd1(x)) / x \|> cmd1() \|> cmd2()
> / >> writeLines() / readr::write_csv()
< readLines()
&& &&
sort \| uniq -c \| sort -nr sort(table(x), decreasing = TRUE)

12.6 Exercises

  1. Run the canonical pipeline from this chapter. Confirm the top line reads 124 rsconnect and the output has exactly 10 lines (check with ... | wc -l).
  2. Which countries download most? Reuse the pipeline but extract field 9 (country) instead of field 7. Confirm the top line reads 1959 NL.
  3. Drop the tail -n +2 stage and re-run. You should see "package" appear in the counts with a count of 1 — proof the header row pollutes results when not skipped.