Get started with tidymedia

library(tidymedia)

tidymedia runs FFmpeg and MediaInfo from R. It helps you prepare media files for research in a way you can repeat. It trims, crops and converts files, often many at once. It also reads media metadata into tibbles.

tidymedia does not try to cover everything FFmpeg can do. The words that FFmpeg uses, such as codec and stream, are defined in the glossary at the end of this page.

This page uses a short sample clip that comes with the package:

video <- system.file("extdata", "sample.mp4", package = "tidymedia")

Start with a task function

Most jobs need one call to a task function. For example, you may need the audio of a recording for a transcription tool. extract_audio() writes the audio to its own file:

extract_audio(video, "audio.m4a")

A task function runs FFmpeg at once. It returns the FFmpeg command it ran, but invisibly, so R prints nothing. A function that writes two files, such as separate_audio_video(), returns both commands.

To see the command without running it, add run = FALSE. The function then returns the command as a string that you can read, log or save:

extract_audio(video, "audio.m4a", run = FALSE)
#> [1] "-y -i \"/private/var/folders/px/frfvbz4n0sx90__c62fwwzz40000gn/T/RtmpI5eZw9/Rinst877f72a946bf/tidymedia/extdata/sample.mp4\" -codec:a copy -vn -map \"0:a:0\" \"audio.m4a\""

This is the main idea of the package. You can read each command before you run it. Cropping works the same way:

crop_video(video, "cropped.mp4", width = 160, height = 120, run = FALSE)
#> [1] "-y -i \"/private/var/folders/px/frfvbz4n0sx90__c62fwwzz40000gn/T/RtmpI5eZw9/Rinst877f72a946bf/tidymedia/extdata/sample.mp4\" -vf \"crop=w=160:h=120:x=(in_w-out_w)/2:y=(in_h-out_h)/2\" -codec:a copy -map \"0:v?\" -map \"0:a?\" \"cropped.mp4\""

Choosing an audio track

Some recordings have more than one audio track, for example a room microphone and a lapel microphone. If you do not choose a track, FFmpeg chooses one for you.

Each task function that reads one input file and picks an audio track has an audio_stream argument. It counts from 0, and it counts only the audio tracks. So audio_stream = 1 is the second audio track, wherever it sits in the file:

extract_audio(video, "lapel.m4a", audio_stream = 1, run = FALSE)
#> [1] "-y -i \"/private/var/folders/px/frfvbz4n0sx90__c62fwwzz40000gn/T/RtmpI5eZw9/Rinst877f72a946bf/tidymedia/extdata/sample.mp4\" -codec:a copy -vn -map \"0:a:1\" \"lapel.m4a\""

If you leave audio_stream out, the default depends on the function. A function that writes exactly one audio track, such as extract_audio(), takes the first track. A function that passes the audio through, such as crop_video(), keeps every track. Compare the -map parts of the two commands above and below:

crop_video(video, "cropped.mp4", width = 160, height = 120, run = FALSE)
#> [1] "-y -i \"/private/var/folders/px/frfvbz4n0sx90__c62fwwzz40000gn/T/RtmpI5eZw9/Rinst877f72a946bf/tidymedia/extdata/sample.mp4\" -vf \"crop=w=160:h=120:x=(in_w-out_w)/2:y=(in_h-out_h)/2\" -codec:a copy -map \"0:v?\" -map \"0:a?\" \"cropped.mp4\""

The help page ?audio_stream lists which functions use each default. It also explains audio_input, which the functions for several input files use. That argument counts input files from 0, not audio tracks.

Each task function has a batch version for a folder of files, such as extract_audio_batch() and crop_video_batch(). See vignette("batch"). For a full research example that uses many task functions, see vignette("workflow").

Three kinds of function

tidymedia has three kinds of function:

The rest of this page shows the pipeline functions.

Building a pipeline

A pipeline starts with ffm_files(), which names the input and output files. You add steps with |>. Each step adds an instruction, and nothing runs yet. ffm_compile() turns the pipeline into the FFmpeg command:

ffm_files(video, "output.mp4") |>
  ffm_trim(start = 1, end = 5) |>
  ffm_crop(width = 160, height = 120) |>
  ffm_codec(video = "libx264") |>
  ffm_compile()
#> [1] "-y -i \"/private/var/folders/px/frfvbz4n0sx90__c62fwwzz40000gn/T/RtmpI5eZw9/Rinst877f72a946bf/tidymedia/extdata/sample.mp4\" -vf \"trim=start=1:end=5,setpts=PTS-STARTPTS,crop=w=160:h=120:x=(in_w-out_w)/2:y=(in_h-out_h)/2\" -codec:v libx264 \"output.mp4\""

When you print a pipeline, R shows the same command. So you can look at a pipeline at any point:

ffm_files(video, "output.mp4") |>
  ffm_scale(width = 320, height = 240) |>
  ffm_pixel_format("yuv420p")
#> tidymedia ffmpeg pipeline:
#> 
#>  -y -i "/private/var/folders/px/frfvbz4n0sx90__c62fwwzz40000gn/T/RtmpI5eZw9/Rinst877f72a946bf/tidymedia/extdata/sample.mp4" -vf "scale=w=320:h=240" -pix_fmt yuv420p "output.mp4"

To run the command and write the output file, use ffm_run() in place of ffm_compile().

More pipeline steps

ffm_fps() changes the frame rate. ffm_drawbox() draws a box on the picture, which can hide a name on screen:

ffm_files(video, "boxed.mp4") |>
  ffm_fps(15) |>
  ffm_drawbox(x = 10, y = 10, width = 60, height = 40, color = "black") |>
  ffm_compile()
#> [1] "-y -i \"/private/var/folders/px/frfvbz4n0sx90__c62fwwzz40000gn/T/RtmpI5eZw9/Rinst877f72a946bf/tidymedia/extdata/sample.mp4\" -vf \"fps=15,drawbox=x=10:y=10:w=60:h=40:c=black:t=fill\" \"boxed.mp4\""

ffm_loudnorm() makes audio a set loudness, in LUFS, with a limit on its true peak. Here ffm_drop() also leaves the video out of the output:

ffm_files(video, "speech.m4a") |>
  ffm_drop("video") |>
  ffm_loudnorm(target_loudness = -23, true_peak = -1) |>
  ffm_compile()
#> [1] "-y -i \"/private/var/folders/px/frfvbz4n0sx90__c62fwwzz40000gn/T/RtmpI5eZw9/Rinst877f72a946bf/tidymedia/extdata/sample.mp4\" -af \"loudnorm=I=-23:TP=-1:LRA=7,asetnsamples=n=4096:p=0\" -vn \"speech.m4a\""

ffm_output_options() adds FFmpeg output options that have no pipeline function of their own. tidymedia still puts them in the right place in the command:

ffm_files(video, "web.mp4") |>
  ffm_output_options("-movflags +faststart") |>
  ffm_compile()
#> [1] "-y -i \"/private/var/folders/px/frfvbz4n0sx90__c62fwwzz40000gn/T/RtmpI5eZw9/Rinst877f72a946bf/tidymedia/extdata/sample.mp4\" -movflags +faststart \"web.mp4\""

Fast cuts and exact cuts

You can cut a clip in two ways. An exact cut re-encodes the video, which is slower. A fast cut uses a stream copy, which keeps the quality but starts at the nearest keyframe.

ffm_seek() does both. Set reencode = FALSE and add ffm_copy() for a fast cut:

# Fast cut with no loss of quality
ffm_files(video, "output.mp4") |>
  ffm_seek(start = 1, end = 5, reencode = FALSE) |>
  ffm_copy() |>
  ffm_compile()
#> [1] "-y -ss 1 -to 5 -i \"/private/var/folders/px/frfvbz4n0sx90__c62fwwzz40000gn/T/RtmpI5eZw9/Rinst877f72a946bf/tidymedia/extdata/sample.mp4\" -codec:v copy -codec:a copy -avoid_negative_ts make_zero -map \"0\" \"output.mp4\""

ffm_seek() uses FFmpeg’s -ss and -to options. ffm_trim() uses FFmpeg’s trim filter. Only ffm_seek() can make a fast cut.

Combining multiple inputs

Some pipeline functions take more than one input. Give ffm_files() a vector of files. Then use ffm_hstack() to put the videos side by side, or ffm_vstack() to put one above the other. ffm_overlay() puts one video on top of another, and ffm_concat() joins them end to end:

ffm_files(c(video, video), "side_by_side.mp4") |>
  ffm_hstack() |>
  ffm_compile()
#> [1] "-y -i \"/private/var/folders/px/frfvbz4n0sx90__c62fwwzz40000gn/T/RtmpI5eZw9/Rinst877f72a946bf/tidymedia/extdata/sample.mp4\" -i \"/private/var/folders/px/frfvbz4n0sx90__c62fwwzz40000gn/T/RtmpI5eZw9/Rinst877f72a946bf/tidymedia/extdata/sample.mp4\" -filter_complex \"[0:v][1:v]hstack=inputs=2:shortest=0[vout]\" -map \"[vout]\" \"side_by_side.mp4\""

ffm_hstack(), ffm_vstack() and ffm_overlay() leave the audio out. To keep it, add ffm_map("0:a"). ffm_concat() keeps all the streams, audio included. One-input task functions that pass audio through, such as crop_video(), do the opposite and keep every audio track.

Two task functions cover the common cases. compare_videos() puts videos side by side or one above the other. picture_in_picture() puts a smaller copy of one video on top of another. Both take audio_input, which names the input whose audio to keep. They copy that audio unchanged unless you choose an audio encoder with audio_codec.

compare_videos(c(video, video), "compare.mp4", audio_input = 0, run = FALSE)
#> [1] "-y -i \"/private/var/folders/px/frfvbz4n0sx90__c62fwwzz40000gn/T/RtmpI5eZw9/Rinst877f72a946bf/tidymedia/extdata/sample.mp4\" -i \"/private/var/folders/px/frfvbz4n0sx90__c62fwwzz40000gn/T/RtmpI5eZw9/Rinst877f72a946bf/tidymedia/extdata/sample.mp4\" -filter_complex \"[0:v][1:v]scale2ref='oh*mdar':'if(lt(main_h,ih),ih,main_h)'[0s][1s];[1s][0s]scale2ref='oh*mdar':'if(lt(main_h,ih),ih,main_h)'[1s][0s];[0s][1s]hstack,setsar=1[vout]\" -codec:a copy -map \"[vout]\" -map \"0:a\" \"compare.mp4\""

A pipeline has one input chain, a list of filters in order, and one output. It cannot build an FFmpeg filter graph with branches. For that, use the direct command ffmpeg():

# Your own arguments, passed to FFmpeg as they are
ffmpeg("-version")[1]
#> [1] "ffmpeg version 9.0.2 Copyright (c) 2000-2026 the FFmpeg developers"

Glossary

Where to next