A 40-minute 4K screen recording took my laptop’s CPU eleven minutes to re-encode just to delete a 90-second dead-air section in the middle. That’s the tax most people pay without realizing there’s a way around it: FFmpeg can cut and splice video in about the time it takes to copy the file, because it never touches a single pixel.

The trick is a feature most FFmpeg tutorials skip past on the way to -vf filters and codec flags: -c copy. Instead of decoding every frame, transforming it, and re-encoding it, stream copy takes the compressed video and audio packets and repackages them into a new container untouched. It’s the difference between re-typing a document to delete a paragraph and just cutting the paragraph out with scissors.


Why This Exists

The scenario is almost always the same. You have a screen recording, a tutorial capture, or a security camera export, and somewhere in the middle there’s a section you don’t want — a pause, a mistake, fifteen minutes of nothing happening. The instinct is to open a video editor, which means a GUI, a project file, a render queue, and a wait proportional to the video’s length and resolution. For a one-off cut on a 4K file, that render can run longer than the video itself.

The Stack Overflow thread that popularized this exact technique has been read enough times to tell you this isn’t a niche problem — it’s the default way people discover FFmpeg can do surgical edits without a timeline.

The catch, and the reason most people don’t reach for this first: stream copy can only cut on keyframes. Video codecs like H.264 don’t store every frame independently — most frames are deltas from the frame before them. You can’t cleanly start a new segment on a delta frame, because it has nothing to be a delta from. That constraint is why this method is fast, and it’s also the one thing that will burn an afternoon if you don’t know about it going in.


The Two-Segment Method

The cleanest way to remove a middle section is to stop thinking of it as “deleting a chunk” and start thinking of it as “keeping two chunks and gluing them together.”

Say you have recording.mp4, it’s 40 minutes long, and you want to remove the section from 12:30 to 14:00.

Step 1 — cut the part before the section you’re removing:

ffmpeg -i recording.mp4 -ss 00:00:00 -to 00:12:30 -c copy part1.mp4

Step 2 — cut the part after it:

ffmpeg -i recording.mp4 -ss 00:14:00 -to 00:40:00 -c copy part2.mp4

Step 3 — write a concat list file:

printf "file 'part1.mp4'\nfile 'part2.mp4'\n" > concat_list.txt

Step 4 — stitch them back together, still without re-encoding:

ffmpeg -f concat -safe 0 -i concat_list.txt -c copy final.mp4

Four commands, and on a 40-minute file this finishes in single-digit seconds — because at no point does FFmpeg decode a frame. It’s reading compressed packets and writing them into a new container, the same way cp reads and writes bytes.

Approach CPU used Time on a 40-min 4K clip Quality loss
Re-encode in a GUI editor Full decode + encode ~11 minutes Yes — every re-encode is generational loss
-c copy two-segment method Packet copy only ~4 seconds None

The Keyframe Problem, and How to Actually Solve It

Here’s where the method breaks for most first attempts: -ss with -c copy snaps to the nearest keyframe, not the exact timestamp you asked for. Ask for 12:30 and FFmpeg will happily cut at 12:27 if that’s where the closest keyframe sits, because that’s the earliest point it can start a valid, decodable stream.

For a screen recording you don’t care about frame-perfect precision on, that’s invisible. For a tutorial where you’re cutting right before someone starts talking, a two-or-three-second drift is exactly the kind of thing a viewer notices and you don’t, because you already know what’s supposed to happen next.

Two ways out, in order of how often you’ll actually need them:

  1. Find your keyframes first, and cut on them deliberately instead of guessing:

    ffprobe -select_streams v -show_frames -show_entries frame=pict_type,pkt_pts_time -of csv recording.mp4 | grep I | head -20
    

    pict_type=I identifies keyframes (I-frames). Round your cut points to the nearest one in that list and the drift disappears — because you’re no longer fighting the codec, you’re working with it.

  2. Re-encode only the tiny window around the cut, and stream-copy everything else. This is the surgical option: decode and re-encode a two-second buffer on each side of your cut point (a fraction of a second of quality loss, invisible at normal viewing distance), while the other 39-something minutes never touch the CPU.

I use option 1 almost every time. Option 2 is for the rare case where the cut point is load-bearing — a demo video where the exact frame someone clicks “deploy” actually matters.

Watch out: -ss placed after -i (as in the examples above) is what you want for this method — it seeks by keyframe on the already-copied stream and stays instant. Placing -ss before -i also seeks by keyframe for -c copy, so either position works here; what won’t work is expecting frame-accurate cuts from either one without re-encoding. The flag position only matters once you introduce re-encoding into the mix.


Turning It Into a Script

The four-command version is fine for a one-off cut. It stops being fine the moment you have to do it eight times for eight files, which is the actual reason I automated it — not because any single cut is hard, but because doing the same four commands by hand eight times in a row is where mistakes creep in.

#!/usr/bin/env bash
# trim_middle.sh — remove [CUT_START, CUT_END] from INPUT, no re-encode
set -euo pipefail

INPUT="$1"; CUT_START="$2"; CUT_END="$3"; OUTPUT="$4"
TMP=$(mktemp -d)
trap 'rm -rf "$TMP"' EXIT

DURATION=$(ffprobe -v error -show_entries format=duration -of csv=p=0 "$INPUT")

ffmpeg -y -i "$INPUT" -ss 00:00:00 -to "$CUT_START" -c copy "$TMP/part1.mp4"
ffmpeg -y -i "$INPUT" -ss "$CUT_END" -to "$DURATION" -c copy "$TMP/part2.mp4"
printf "file '%s'\nfile '%s'\n" "$TMP/part1.mp4" "$TMP/part2.mp4" > "$TMP/list.txt"
ffmpeg -y -f concat -safe 0 -i "$TMP/list.txt" -c copy "$OUTPUT"

echo "Done: $OUTPUT"

Run it as ./trim_middle.sh recording.mp4 00:12:30 00:14:00 final.mp4 and it does all four steps, cleans up its temp files, and tells you when it’s done. Point a for loop at a folder of recordings with the same dead-air pattern — a fixed intro slate, a recurring pause before a demo starts — and eight manual edits become one command.


When Stream Copy Isn’t the Right Tool

This method has a real ceiling, and pretending otherwise is how people end up filing bug reports against FFmpeg for “corrupting” their video. Stream copy will not help you if:

  • You need frame-accurate cuts on every single edit and can’t tolerate rounding to the nearest keyframe. Re-encoding is not optional there.
  • You’re also changing resolution, codec, or container in a way that isn’t just remuxing — -c copy only works when the two segments you’re concatenating share the same codec parameters end to end.
  • The source has variable frame rate or unusual GOP structure — screen recorders and some phone cameras produce erratic keyframe spacing, and the concat demuxer can produce sync drift between audio and video in the joined file. Test the seam before trusting a batch of them.

For everything else — screen recordings, webinar exports, tutorial captures, security footage — the two-segment method turns a coffee-break render into something you don’t have time to get coffee during.