How Smoovie Makes 24fps Look Like 60: Inside Our Motion Interpolation Engine
The problem
Almost every film and TV show you watch was shot at 24 frames per second. Your phone's screen refreshes 60, 90 or 120 times a second. So for most of the time your display is awake, it is showing you a frame that is already stale — the same picture held on screen for two, three, sometimes five refreshes in a row. When the camera pans, your eye tracks the motion smoothly but the image jumps in discrete steps. That stepping is judder, and once you notice it you can't un-notice it.
The fix is conceptually simple and practically brutal: invent the missing frames. If you have frame A at 0ms and frame B at 41ms, generate a plausible frame for 20ms in between. Do it fast enough to keep up with playback, and do it well enough that nobody sees the seams.
This is called MEMC — Motion Estimation, Motion Compensation. Smoovie has its own implementation, built from scratch, running entirely on-device. This post is a tour of it.
Why this is hard on a phone
The reference implementation everyone knows is SVP. It's excellent, and it's a desktop tool. Its Android build needs a Snapdragon 865-class chip just to manage 1080p, and its architecture explains why: motion estimation runs on the CPU (a port of MVTools with hand-written NEON kernels), frame synthesis runs on the GPU via Vulkan compute.
Our first instinct was to beat that by moving everything to the GPU. That was wrong, and finding out why shaped the whole engine.
We built GPU motion estimation. It worked. It was also fighting the same GPU that has to composite and present the frame, and on a phone that's one chip with one memory bus. Push the compute harder and the display path starves. We measured playback that reported ~80fps in the HUD while actually running at about 10fps wall-clock, with ~92% of frames dropped at the output. The GPU wasn't slow — it was contended.
So the shipping architecture converged on the same split SVP uses, for the same reason, and then went further:
The CPU finds the motion. The GPU draws the frames. They run at the same time, one frame-pair apart.
The CPU cores are idle anyway during video playback (hardware decode barely touches them). The GPU is the scarce resource. Handing motion search to the CPU doesn't just parallelise the work — it takes ~5ms of pressure off the exact chip that also has to get pixels to the panel.
The pipeline
decode -> vf_smoovie -> [1] CPU: motion estimation (big cores, NEON, ~5ms)
[2] CPU: field regularization
[3] GPU: warp + colour (Vulkan compute, ~11ms)
[4] GPU: paced present (our own swapchain)
Stages 1–2 run on one thread, 3 on another, 4 on a third — all overlapping. While the GPU synthesises pair N, the CPU is already estimating motion for pair N+1. Throughput is the max of the two stages, not the sum. (mpv itself runs headless: we replaced its video output entirely, so nothing queues frames behind our back.)
At 1080p on a Snapdragon 865, per frame pair:
- CPU motion estimation — 5.3 ms
- GPU synthesis — 11.1 ms
- Budget at a 30fps source — 33.3 ms
About two-thirds of the budget is idle. That headroom is the whole point — it's what keeps the display path from starving.
Stage 1 — Finding the motion
In plain English: chop the frame into a grid of squares, and for each square, search the next frame for where that square went.
Each frame is divided into 32×32 pixel blocks. For every block we search frame B for the position that best matches it, scored by SAD (sum of absolute differences — literally, how different are these two squares, pixel by pixel). The winner's offset is that block's motion vector.
Done naively this is catastrophically slow. A ±32 pixel search is 4,225 candidate positions × 1,024 pixels each × ~2,500 blocks. Our first CPU implementation took 191ms per frame pair. Real-time needed under 33. Here's how it got to 5:
A pyramid. We build half-resolution and quarter-resolution copies of both frames. The search starts on the smallest one, where a 32-pixel motion is only 8 pixels, then refines downward. Three levels. Big motions get found by a cheap coarse search rather than an expensive fine one.
Diamond search, not exhaustive search. From a starting guess, we test a small diamond of positions around it, jump to the best one, and repeat until the centre wins — then one fine refinement pass. That's 10–20 SAD evaluations instead of thousands. The catch is that it's only as good as its starting guess, which is why we feed it predictors: the same block's vector from the previous frame, the neighbouring block's vector, and a global pan estimate (the median of last frame's entire field — median rather than average, so one fast-moving object can't drag the camera estimate off).
NEON, everywhere. The inner SAD loop compiles to ARM's UABD/UADALP instructions — 16 pixels of absolute-difference-and-accumulate per instruction. The accumulator is held across four rows before the expensive cross-lane reduction, and there's an early-out: the moment a candidate's running SAD exceeds the current best, we stop summing.
Pinned to the big cores. This one is the single largest speedup and the least obvious. Modern phone chips are big.LITTLE — on a Snapdragon 865, four weak efficiency cores and four fast ones. Spreading the work across all eight means every parallel step waits on the slowest core:
- Naive scalar, all 8 cores — 191 ms
- NEON + padded planes, all 8 cores — 36 ms
- NEON + padded planes, 4 big cores only — 15.6 ms
Same work, 2.3× faster, purely from not using half the cores. In-app, with pyramid reuse and other trims, it now runs at ~5ms.
Quarter-pixel precision. Real motion isn't a whole number of pixels. After the integer search finds a winner, we evaluate the SAD one pixel either side and fit a parabola through the three points — the minimum of that curve is the sub-pixel offset, in quarter-pixel steps. Four extra SAD evaluations per block over memory the search just read, essentially free. Blocks that already match near-perfectly skip it (their error curve is flat, so fitting a parabola to it just fits noise).
Stage 2 — Cleaning up the vector field
In plain English: the raw search result is noisy. A block of blank sky or repeating fence can "match" almost anywhere. Fix the outliers before drawing.
A minimum-SAD winner is not always the true motion. Featureless regions, periodic textures (railings, window grids, brickwork) and noise all produce confident-looking nonsense. If you warp with a noisy field you get pixels swimming around inside otherwise still objects.
Three passes, all on the CPU, on the same thread that just did the search:
1. 3×3 vector median. Each block's vector is replaced by the median of itself and its eight neighbours (x and y independently). Outliers vanish; genuine motion boundaries survive, because a median doesn't smear across an edge the way an average does.
2. Global-pan pull. We compute a confidence-weighted average motion for the whole frame. Any block whose SAD says it matched badly gets pulled toward that global vector, proportionally to how untrustworthy it is. A block that found nothing convincing follows the camera, which is almost always right.
3. Temporal hysteresis. We keep a damped copy of the previous frame's final field. A block oscillating between two nearly-equal candidates settles onto the average of the two — which is usually the true fractional motion the pixel grid couldn't express in the first place.
The tuning here (70% global pull, error threshold 40, coherence penalty 4 per pixel of deviation) is deliberately close to SVP's "truemotion" preset. Those constants exist because SVP's authors spent years finding them.
Stage 3 — Drawing the frame
In plain English: for every pixel in the new frame, look backward along the motion into frame A and forward into frame B, and mix.
This is a Vulkan compute shader, one thread per four output pixels. For an output frame at time t between A and B, each pixel does:
va = sampleA(x - t*mx, y - t*my); // trace back into A
vb = sampleB(x + (1-t)*mx, y + (1-t)*my); // trace forward into B
out = mix(va, vb, t);
Four ideas sit on top of that core, and each one exists because of a specific visible artifact:
OBMC — overlapped block motion compensation. If every pixel simply used its own block's vector, you'd see the block grid as a lattice of hard tears across moving objects. Instead each pixel bilinearly blends the four surrounding block centre vectors. Motion becomes a smooth field instead of a staircase.
A soft occlusion mask. When an object moves, it uncovers background that exists in only one of the two frames. There's nothing to match, and forcing a match produces smear. So we ramp: where the forward and backward samples strongly disagree, or where the block's match quality was poor, the output fades toward a plain cross-dissolve. Bad motion degrades into softness rather than into garbage. The thresholds matter more than they look — an earlier, tighter version fired almost everywhere on real footage (sensor noise alone crossed it) and quietly dissolved the entire frame back toward 24fps. Trusting the motion more, and only dissolving at genuine occlusions, is what actually produced smoothness.
Motion-compensated temporal denoise. Here's a free win. The two samples are motion-*aligned* observations of the same point in the scene. Where they agree, the difference between them is independent sensor and compression noise — so biasing the blend toward 50/50 averages that noise down by ~30%, using taps we already fetched. We gate it to slow motion and confident matches, because averaging two mis-aligned taps is how you get the "jelly" wobble.
Byte-packing. Each shader invocation computes four horizontal pixels and writes them as a single 32-bit word. This started as a bandwidth fix during bring-up, and it was a big one: reading results back went from 329ms to 4ms.
Luma and chroma are warped by separate specialised shaders, then a final pass converts YUV to RGB, applies the per-video colour matrix and (for HDR content) PQ or HLG tone mapping — and writes directly into the presentable image, skipping an intermediate copy worth about 2ms per frame.
Stage 4 — Pacing
In plain English: generating perfect frames isn't enough. You have to show them at perfectly even intervals, or it still looks like judder.
This one surprised us. We had good frames and the motion still looked "steppy" — identically on 60Hz and 120Hz panels, which was the clue. It wasn't a quality problem at all.
The synthesis stage is bursty: some frame pairs need two output frames, some need three. Presenting each one the moment it's ready meant one frame held for one refresh, the next for four. Perfectly good frames, terrible rhythm.
The fix is a proper pacing stack:
- Pin the panel to its maximum refresh rate. Never lower it. A 60fps stream on a 120Hz panel should hold each frame exactly 2 refreshes — that's smoother than switching the panel to 60Hz.
- Schedule presents explicitly. Using VK_GOOGLE_display_timing, every frame carries an exact target present time, snapped to a whole number of refresh intervals.
- Phase-lock to the compositor. The driver reports when frames actually appeared. We use that to anchor our schedule, so it can't drift from the system compositor's clock.
- Sticky hold counts. We measure the input cadence over half-second windows and only adopt a new hold count after it persists — otherwise the schedule flaps between 1 and 2 refreshes and you see the frame rate wobble.
Because we present through our own Vulkan swapchain, we own this end-to-end. Nothing else is queuing frames behind our back.
The cases that need special handling
Scene cuts. Interpolating across a hard cut means morphing one shot into an unrelated one — a wet, horrible smear on every edit. We detect cuts from data the motion search already produced (a spike in whole-frame match error, plus the fraction of blocks that failed to match), and on a cut the shader stops interpolating and simply repeats the nearer source frame. A cut should stay a cut. Detection uses two independent triggers because the obvious one is blind to dark scenes and gets capped by letterbox bars — black bars match perfectly across any edit, and they drag the average down.
High-frame-rate sources. Content already at 60fps or above gets nothing from interpolation, and worse, it floods a pipeline built around a 60fps output grid. A cadence estimator detects it and switches to clean pass-through, presenting the frames a normal player would.
Slower devices. There's a 30fps tier that subsamples the same 60fps grid, halving both the synthesis and present work. It's an evenly-spaced grid, not alternate pairs — that distinction matters, and getting it wrong produces uneven 1/60, 3/60 spacing that looks worse than doing nothing.
Things we measured that we would have guessed wrong
Almost every intuition about GPU performance turned out to be false on Adreno hardware. Some favourites:
- Half-precision math: zero effect. The warp shader isn't limited by arithmetic, it's limited by texture taps. Halving the precision of maths that isn't the bottleneck buys nothing.
- Caching motion vectors in registers: 6% slower. Fetching each thread's block vectors once instead of per-pixel seemed obviously good. The vector field is tiny and already sits in L2 cache, so the reads were free — and the register usage cut GPU occupancy. Reverted.
- Dead code costs 23%. A branch that is never taken still burdens the shader, through register pressure and scheduling. We now compile separate shader variants rather than leaving disabled features behind runtime flags. Several quality features that validated correctly on-device were deleted anyway, purely because their presence cost more than their absence gained.
- Fixing the wrong thing feels productive. The "everything looks steppy" complaint sat in the quality bucket for weeks. It was pacing.
What it still can't do
Honesty is more useful than a highlight reel.
Small fast objects ghost. A football is about 30 pixels across; blocks are 32. Any block containing the ball is mostly crowd, so minimum-SAD picks the crowd's motion and the ball never gets a vector of its own. Untracked, it gets drawn twice — once at each source position. Tuning cannot fix this; it needs sub-block motion search or a separate object layer. We tested the obvious knobs and confirmed they do nothing: the search's greedy descent never even evaluates the ball's true vector, because every step toward it is a worse match on the crowd behind it.
4K is over budget. At 4K the synthesis costs ~76ms per pair against a 41.7ms budget. The fix is known and not yet built: motion-estimate and warp at panel resolution rather than source resolution. We currently process 8.3 megapixels to display 2.1 — and then throw the difference away in the final scaling blit. That's a 4× win sitting on the table.
The motion field is anchored to the first frame. Vectors are indexed by output position but were measured on frame A's grid, so under very large motion the vector that belongs at a pixel actually lives some distance away. Proper solutions maintain separate forward and backward fields.
Try it
The engine ships in Smoovie with three settings: off, 30fps and 60fps. There is deliberately no 90 or 120 target — beyond 60 the artifacts grow faster than the smoothness does, and every extra output frame is real thermal budget on a device you're holding in your hand.
If you've ever watched a pan across a landscape and felt the image stutter against motion your eye was tracking perfectly smoothly — turn it on. That's the thing it fixes.
Comments
Great work! Thank you for this.
Leave a comment
Comments are reviewed before they appear.