Eulerian video magnification reveals temporal variations in videos that are difficult or impossible to see with the naked eye. It can amplify subtle color changes — like the flush of blood under skin with each heartbeat — or tiny motions, making them clearly visible.
Figure 1: Original video, 20X magnified, and 100X magnified.
This is a Python implementation of MIT CSAIL's paper, "Eulerian Video Magnification for Revealing Subtle Changes in the World" (Wu et al., SIGGRAPH 2012).
There are two fundamental approaches to analyzing motion in video:
- Lagrangian — track individual points across frames (optical flow). Works well for large motions but struggles with sub-pixel changes.
- Eulerian — observe how pixel values change over time at fixed spatial locations. This is what EVM uses.
The key insight is that for small motions, the temporal intensity change at a fixed pixel is proportional to the spatial gradient multiplied by the displacement. By amplifying these temporal changes, we can make invisible motions visible — without ever computing motion trajectories.
The algorithm has four main stages:
1. Spatial Decomposition (Laplacian Pyramid)
Each video frame is decomposed into a Laplacian pyramid — a multi-scale representation where each level captures spatial details at a different frequency band. This separates fine details from coarse structure, allowing the algorithm to amplify motion at specific spatial scales independently.
A Gaussian pyramid is built by repeatedly downsampling with cv2.pyrDown. The Laplacian pyramid is the difference between consecutive Gaussian levels:
2. Temporal Filtering (Bandpass)
At each spatial location and pyramid level, pixel values are treated as a 1D time-series signal. An ideal bandpass filter (implemented via FFT) extracts only the temporal frequencies of interest:
- For color magnification (e.g., pulse detection): low frequencies, typically 0.5–2 Hz
- For motion magnification (e.g., vibrations): higher frequencies matching the motion
The FFT is computed along the time axis, frequencies outside
3. Amplification
The filtered signal is multiplied by an amplification factor
where
4. Reconstruction
The modified Laplacian pyramid is collapsed back into a full-resolution video by iteratively upsampling and adding:
The theoretical justification for why temporal filtering reveals motion comes from a first-order Taylor expansion. For a 1D image signal
By Taylor expansion:
The temporal variation at a fixed pixel
The motion
| Application | Frequency Band | Amplification | What It Reveals |
|---|---|---|---|
| Pulse detection | 0.5–2 Hz | 50–150x | Blood flow causing subtle skin color changes |
| Breathing | 0.1–0.5 Hz | 10–30x | Chest/body movement during respiration |
| Structural vibration | 1–50 Hz | 20–100x | Building sway, bridge vibrations |
| Musical vibration | 50–500 Hz | 50–200x | Object vibrations from sound |
-
Artifacts at high amplification — when
$\alpha$ is too large relative to the spatial wavelength, the first-order approximation breaks down and produces ringing/ghosting artifacts. - Noise amplification — the algorithm amplifies all temporal variations in the frequency band, including sensor noise. Low-light or noisy videos produce poor results.
- Large motion — the Eulerian approach assumes small motions. Objects with significant displacement across frames will not be correctly magnified.
- No occlusion handling — since we observe fixed pixel locations, occluded regions cannot be recovered.
Video is converted to YIQ (NTSC) color space using the same matrices as MATLAB's rgb2ntsc/ntsc2rgb. This separates luminance (Y) from chrominance (I, Q), enabling independent control of color amplification via the --chrom-attenuation flag.
Ideal bandpass filtering via FFT: in-band frequencies pass at full amplitude and everything else is removed, and the filtered signal keeps its sign, so pixels oscillate above and below their mean. The amplified output is therefore (1 + α) × the in-band change on every amplified level. (MATLAB's ideal_bandpassing.m keeps positive frequencies only, which halves the in-band signal, so there α delivered about α / 2; this version doesn't, since v3.0.0.)
Before the FFT the clip is extended with its time-reversed copy. The FFT treats a clip as a loop, so without this the last frame is joined to the first, and slow drift over the clip (lighting, auto-exposure) turns into amplified flicker. The band must contain at least one FFT bin; a band narrower than the clip's resolution (fps / frames) gives a warning, and a band above the Nyquist frequency (fps / 2) is an error.
Per-level alpha is computed based on lambda_c and the representative spatial wavelength at each pyramid level (Figure 6 of the paper). This prevents over-amplification of fine spatial details beyond what the first-order Taylor expansion supports. Per-level alpha is never negative (the reference formula goes negative for small frames, which would shrink the motion). Two levels are zeroed out:
- Level 0 (finest, full resolution) — captures the highest spatial frequencies (sharpest edges and fine details). The spatial wavelengths are so short that even modest amplification breaks the first-order Taylor approximation, producing ringing and ghosting artifacts.
- Coarsest level (low-pass residual) — this is not a true bandpass level; it is the Gaussian remainder (
gauss[-1]) appended directly to the pyramid. It contains the DC component (overall mean intensity), so amplifying it would shift global brightness rather than reveal temporal variations.
Collapsing a Laplacian pyramid is linear, and only the levels with a non-zero gain change, so the output is input + collapse(amplified levels). The pipeline therefore stores only the amplified levels (1 to N-2, at most a third of the video's size) and never level 0 or the low-pass residual. The CLI keeps the decoded frames as uint8, converts them to YIQ a block of 8 frames at a time, and writes each output block as soon as it's ready. The temporal filter works on 32 MB chunks of pixels.
Measured on face.mp4 (301 frames, 528×592, 4 levels, -a 50 -fl 0.83 -fh 1.0 --lambda-c 1000) on an idle Ryzen 7 7445HS (12 threads) with an RTX 4050 Laptop 6 GB; times are the median of 5 runs after a warm-up, including decoding and writing:
.avi (MJPG) |
.mkv (FFV1, lossless) |
peak memory | |
|---|---|---|---|
| CPU | 4.5 s | 7.4 s | 0.91 GiB RAM (4.1 GiB in v2.1.0) |
GPU (--gpu) |
1.5 s | 4.6 s | 0.88 GiB VRAM |
Lossless output costs about 10 ms a frame: OpenCV's FFV1 encoder runs on one core.
As a rule of thumb, RAM is about 2.3× the uint8 video (frames × height × width × 3 bytes) plus 0.4 GB. The GPU needs about 1.25× the uint8 video in VRAM plus 0.4 GB; --gpu prints its estimate before starting.
- Nyquist and empty-band validation, progress reporting with ETA
- Synthetic validation against exact ground truth (see below)
The easiest way to try the notebook — click the badge at the top of this README. No installation needed.
CLI tool (recommended for processing your own videos):
git clone https://github.com/joeljose/Eulerian-Video-Magnification.git
cd Eulerian-Video-Magnification
pip install -r requirements.txt
python evm.py -i input.mp4Notebook (for interactive exploration and learning):
pip install -r requirements.txt matplotlib requests
jupyter notebook Eulerian_Video_Magnification.ipynbRequirements: Python 3.10+ (tested on 3.10 and 3.11)
# Build (tags as evm:<version> and evm:latest)
./docker-build.sh
# Run (--user lets the container write files you own into the mounted folder)
docker run --rm -it --user "$(id -u):$(id -g)" \
-v "$(pwd)":/app/data \
evm \
-i /app/data/input.mp4 -o /app/data/output.avidocker build . needs no build arguments. The image runs as a fixed non-root user and installs hash-locked dependencies from requirements.lock.
For faster processing on NVIDIA GPUs.
Prerequisites:
- NVIDIA GPU with a driver that supports CUDA 12 (driver 525 or newer)
- NVIDIA drivers installed on the host
- Docker
- NVIDIA Container Toolkit — allows Docker to access the GPU
Verify your GPU is accessible:
nvidia-smi # Should show your GPU name, driver version, and CUDA versionBuild the CUDA Docker image:
./docker-build-cuda.shRun on your video:
# Basic usage — magnify face.mp4 with default settings
docker run --gpus all --rm --user "$(id -u):$(id -g)" \
-v "$(pwd)":/data \
evm-cuda \
-i /data/face.mp4 -o /data/face_magnified.avi
# Pulse detection (0.83–1.0 Hz, coarse levels only)
docker run --gpus all --rm --user "$(id -u):$(id -g)" \
-v "$(pwd)":/data \
evm-cuda \
-i /data/face.mp4 -o /data/face_magnified.avi \
-fl 0.83 -fh 1.0 -a 50 --lambda-c 1000
# Select a specific GPU (for multi-GPU systems)
docker run --gpus all --rm --user "$(id -u):$(id -g)" \
-v "$(pwd)":/data \
evm-cuda \
-i /data/input.mp4 -o /data/output.avi --device 1The image is Python 3.11 slim plus CuPy and the CUDA libraries as pip wheels (about 3.5 GB; no CUDA base image), and runs evm.py --gpu: the same code as the CPU version, on CuPy arrays (backed by cuFFT). The CPU and GPU give the same result to float precision. It automatically checks available VRAM before processing and exits with a clear error if the video is too large.
VRAM requirements: Depends on video resolution and length; the tool prints its estimate before starting (measured to be within 25%). face.mp4 (301 frames, 528×592) needs about 0.9 GB; a 1080p 30 s clip at 30 fps needs about 7.6 GB of VRAM, plus about 5.6 GB of host RAM for the decoded frames. See Memory and Speed.
python evm.py -i face.mp4
python evm.py -i face.mp4 -o pulse.avi -a 50 -fl 0.83 -fh 1.0 --lambda-c 1000
python evm.py -i guitar.mp4 -fl 72 -fh 92 -a 50 --lambda-c 10 --chrom-attenuation 0| Flag | Default | Description |
|---|---|---|
-i / --input |
(required) | Input video path |
-o / --output |
<input>_magnified.avi |
Output video path. The extension picks the format: .mkv (FFV1, lossless), .avi (MJPG, lossy) or .mp4 (MPEG-4, lossy). Use .mkv for analysis: JPEG compression noise can be as large as the sub-1% colour changes EVM reveals |
-fl / --freq-low |
0.5 | Lower cutoff frequency (Hz) |
-fh / --freq-high |
2.0 | Upper cutoff frequency (Hz) |
-a / --amplification |
10 | Amplification factor α: in-band changes on each amplified level are multiplied by exactly 1 + α (capped per level by --lambda-c) |
--pyramid-levels |
4 | Number of Laplacian pyramid levels |
--lambda-c |
16 | Cutoff spatial wavelength in pixels (paper Figure 6). Structures smaller than this get reduced amplification, so lower = stronger amplification. The effective per-level gains are printed at startup. |
--chrom-attenuation |
1.0 | Color channel attenuation (0=luminance only, 1=full) |
--fps |
(from video) | Frame rate of the input, for files that don't report one |
--keep-audio |
off | Copy the input's audio track into the output. Needs ffmpeg on the PATH (not included in the Docker images); if the container can't hold the audio codec, it warns and keeps the video without audio |
--version |
— | Show program version and exit |
Add --gpu to run on an NVIDIA GPU. The recommended way is via Docker (see GPU setup above). If running outside Docker:
pip install -r requirements-cuda.txt # or: pip install .[cuda]
python evm.py -i face.mp4 --gpu
python evm.py -i face.mp4 -a 50 -fl 0.83 -fh 1.0 --lambda-c 1000 --gpu --device 0| GPU Flag | Default | Description |
|---|---|---|
--gpu |
off | Run on an NVIDIA GPU with CuPy |
--device |
0 | CUDA device ID (for multi-GPU systems) |
The tool prints the GPU name, estimated VRAM usage, and available memory before processing. python evm_cuda.py … still works for one release; it runs evm.py --gpu.
pip install . (or pip install .[cuda] for the GPU) installs an evm command, and import evm gives the functions (load_video, eulerian_magnification, save_video, …). The pipeline functions accept NumPy or CuPy arrays.
Open the notebook and run all cells. On Colab it clones this repository and imports evm.py, so it runs the same code as the CLI, on the bundled face.mp4 from the original paper. To use your own video, upload it and change the filename variable.
- Use
show_frequencies()in the notebook to visualize frequency content before choosing cutoff frequencies. - Start with low amplification and increase gradually.
- For pulse/color magnification: a narrow band around the heart rate (e.g. 0.83–1.0 Hz), high α, and a large
--lambda-c(e.g. 1000) so the fine levels, which mostly carry sensor noise, are amplified less. Check the printed level gains. - For motion magnification: match the frequency band to the motion you want to reveal, and keep the magnified motion under about 1 px (see below).
At startup the tool prints the effective band and the gain applied to each pyramid level. Two things decide how much you actually get:
-
Per-level gains. In-band changes on a level with gain
gare multiplied by exactly1 + g. Level 0 (finest) and the coarsest level (the low-pass residual) are never amplified. The levels in between getα, capped by--lambda-c, which shrinks the gain of levels whose spatial wavelength is short compared withλc. For face.mp4 (528×592, 4 levels):--lambda-cgains at -a 10(finest → coarsest)gains at -a 5016 (default) 0, 10, 10, 0 0, 50, 50, 0 80 0, 10, 10, 0 0, 50, 50, 0 1000 0, 0, 0.9, 0 0, 4.7, 11.5, 0 -
Spatial content. Only the spatial detail carried by the amplified levels is magnified. On synthetic shapes a thin ring moves about 0.57× the requested amount and a filled circle about 0.44×, because part of their edges sits in level 0 or the low-pass residual (see Synthetic validation).
-
Frequency resolution. A clip of
nframes atfpscan only separate frequenciesfps / napart: 0.1 Hz for face.mp4 (301 frames at 30 fps). A narrower band gives a warning, and a band that contains no frequency bin at all is an error. Use a longer clip for narrow bands.
scripts/synthetic_shapes.py renders circles, squares and rings whose size pulses by an exact sub-pixel amount, magnifies them, and measures the result against the ideal. Run it with python scripts/synthetic_shapes.py; tests/test_synthetic_shapes.py keeps the key results as tests. Findings, asking for 10× (-a 9, since the output is 1 + α times the input), band 0.5–3 Hz, --lambda-c 10, 4 levels, on 128×128 frames:
| Check | Result |
|---|---|
| Motion inside the band | about 0.57× the requested gain on a thin ring, 0.44× on a filled circle: the finest level and the low-pass residual, which carry the rest of the edge, are not amplified |
| Motion outside the band | unchanged (1.0×) |
| Phase lag | none |
| Circles vs squares | same gain in every direction (about 3% spread) |
| Large motion | quality drops quickly once the magnified motion passes about 1 px; ghost edges dominate from 2 px |
The detailed numbers are in docs/research/synthetic-validation.md.
All tests run inside Docker — no local Python dependencies needed. Build the test image once, then run tests as many times as you need:
# Build the CPU test image (first time, or after Dockerfile/dependency changes)
./docker-build.sh
# Run CPU unit tests (builds image automatically if not found)
./test.sh
# Run GPU unit tests (requires NVIDIA GPU + Container Toolkit)
./test.sh gpu
# Force rebuild before testing
./test.sh --buildCPU tests (tests/test_evm.py) cover:
- Color conversion roundtrip (YIQ ↔ RGB)
- Bandpass filter (passband, rejection, DC)
- Laplacian pyramid (reconstruction roundtrip, shapes, finite values)
load_videoerror handling (unreadable input, fps 0, under-reported frame count)save_videofailing loudly when the writer cannot open- Empty frequency bands and per-level gains
- All CLI input validation error paths, and one CLI run to completion on a real clip
- A golden regression test: a 48-frame crop of face.mp4 and its stored output (
tests/data/golden_face.npz). Regenerate it withpython scripts/make_golden.pyonly when an output change is intended, and say so in the CHANGELOG
CUDA shim tests (tests/test_evm_cuda_shim.py) run the GPU branch of evm.py on the CPU with a fake cupy (NumPy/SciPy underneath), so the GPU code path, --gpu setup and VRAM check are tested without a GPU; the result must equal the CPU result.
GPU tests (tests/test_evm_cuda.py, real CuPy) cover:
- GPU color conversion roundtrip
- GPU pyramid operations against OpenCV (to 1e-5, odd sizes too)
- GPU bandpass filter
- CPU vs GPU end to end (PSNR ≥ 60 dB)
Dev workflow:
- Make your changes
- Run
./test.sh(or./test.sh gpufor CUDA changes) - If all tests pass, commit and open a PR
- CI runs lint, the unit tests, and a full pipeline run that checks the output is magnified
Version is tracked in a VERSION file at the project root, and evm.py has __version__ baked into the source (updated at release time; a test checks they match). pyproject.toml reads the version from evm.py.
To cut a release:
- Update
VERSIONwith the new version number - Update
__version__inevm.py(e.g.,"3.0.0") - Update
CHANGELOG.md— move items from[Unreleased]to[X.Y.Z] - YYYY-MM-DD - Commit:
Release vX.Y.Z - Tag:
git tag -a vX.Y.Z -m "Release vX.Y.Z" - Push:
git push && git push origin vX.Y.Z - Rebuild Docker images:
./docker-build.shand./docker-build-cuda.sh
Docker build scripts read from VERSION and tag images accordingly (e.g., evm:3.0.0, evm-cuda:3.0.0). Images also carry a version label visible via docker inspect.
evm.py # CLI tool and library (CPU, or GPU with --gpu)
evm_cuda.py # Deprecated: runs evm.py --gpu
Dockerfile # CPU Docker image
Dockerfile.cuda # GPU Docker image
docker-build.sh # Build + tag CPU image
docker-build-cuda.sh # Build + tag GPU image
test.sh # Run unit tests (cpu/gpu)
requirements.txt # CPU runtime dependencies
requirements-cuda.txt # GPU runtime dependencies
requirements-dev.txt # Dev dependencies (pytest, ruff), pinned exactly
pyproject.toml # Packaging (pip install .) and ruff configuration
tests/
test_evm.py # CPU unit tests
test_evm_cuda.py # GPU tests (real CuPy)
test_evm_cuda_shim.py # GPU code path on CPU with a fake cupy (no GPU needed)
test_synthetic_shapes.py # Gain, lag and isotropy on shapes with exact ground truth
scripts/
synthetic_shapes.py # Synthetic pulsating shapes and measurements
make_golden.py # Regenerates tests/data/golden_face.npz
docs/design/ # Architecture decision records
VERSION # Single source of truth for version
CHANGELOG.md # Release history
-
Wu, H-Y., Rubinstein, M., Shih, E., Guttag, J., Durand, F., & Freeman, W. (2012). Eulerian Video Magnification for Revealing Subtle Changes in the World. ACM Transactions on Graphics (SIGGRAPH), 31(4).



