Article Depth map scroll
Depth map scroll
A component that stacks image layers with their own depth maps on top of each other, opened and lit as you scroll. How it works, from the maps to the candle.
24 min read · Part of Depth map scroll, where the component itself is live
What you are looking at
Depth map scroll is one section of a web page: a small label, a large heading, a framed painting and a caption. The painting is the point. It is four cut-out image layers, each with its own depth map, stacked back to front, so when the view shifts the figure turns in space instead of sliding like a flat card.
Scroll inside the frame below. The rounded window grows, the scene zooms and gains depth, and the hand holding the glass reaches out past the frame. Move the pointer over the painting: the view follows it, and a warm candle light brushes across the paint.
Depth map scroll, the live component
Live, try itFour things happen, and the strip below shows them live as you scroll the component above. The line that moves is where you are in the scroll.
Each section from here opens with a plain answer: what the part does, and where it is not obvious, why it exists or what you would see without it. Read only those and you will understand the component. The text and the interactive figures underneath go one level deeper each, down to the code.
Peel the painting apart
- What it does
- The painting is not one image. It is four cut-out layers stacked back to front: the sky, the carved building, the woman (the layer is called model), and the glass in her hand.
- Why it exists
- Things can only move against each other if they are separate pieces. Cutting the painting apart is what lets the woman shift in front of the arch.
Below, the four layers are real planes in 3D. Move the pointer from left to right and the stack turns; far enough round and you see the layers edge on, from the side.
Peel the painting apart
InteractiveHide the building and the model sits in front of the sky with nothing in between; hide the model and the arch shows the see-through hole the sky appears in. The layers are drawn in a fixed order, sky first and glass last, each one painted over the one before.
What is a depth map?
- What it does
- A depth map is a black-and-white picture that says how near each part of an image is. Brighter is nearer, black is far.
- Why it exists
- A flat picture has no idea what is in front of what. The depth map is how the component knows the rim of the glass is closer to you than the hand behind it.
On its own a depth map looks like a soft photograph of the shape. Read as height, it lifts a flat picture into a surface. Turn the surface below and you are looking at exactly what this component works with.
Lift the painting by its depth map
InteractiveEvery layer ships with its own map, made by a depth model (section 15), so the component loads eight images: four colour layers and four depth maps, all 2200×1736. The glass is the clearest: a rounded cylinder whose rim comes forward. The woman is subtler, because faces were the hardest thing for the depth model, and her map was assembled by hand from separate head and body passes.
One level deeper. Turn off flood the edges and look at the glass from the side. Outside the cut-out the depth drops to zero, so the surface falls off a cliff at the edge and stretches into walls of smeared pixels. The fix is done when the images are prepared: each map is flooded past its outline, filled with depth that continues smoothly from the edge. Nothing out there is ever drawn, but the surface has no cliff to fall off.
Shape is not placement
- What it does
- Each layer is given its own slot on the scale from far to near. The slot is called its band.
- Without it
- All four layers would sit in the same space and cut through each other. The sky would bulge as far forward as the glass.
Switch the figure to raw maps to see the problem, then back to placed in bands to see the fix. The picture is the four real layers seen from the side; the bars underneath are the same thing as a ruler.
Give every layer its own slot
InteractivePlaced. Each layer keeps its shape but lives in its own slot, so the sky stays behind and the glass stays in front.
One level deeper. A band is two numbers, where the slot starts and where it ends, and the placement is one line: placed = lo + raw · (hi − lo). The slot's width is a design choice. A narrow band makes a layer read as a flat card far away (the sky gets 0.00 to 0.10); a wide band gives it deep relief (the model gets 0.20 to 0.55). Drag model's band width to feel the difference.
One thing in the figure looks wrong and is not. The building's band (0.60 to 0.85) is nearer than the model's, although she is drawn on top of it. Draw order decides what covers what; the band decides how much a layer moves. The model's band sits around the one depth that never moves (0.36, the focus), so she stays almost still while the arch shifts around her, like a doorway you are looking through.
A mesh, not a card
- What it does
- Each layer is drawn on a fine net of points, 160 squares across and 160 down. The depth map pushes every point nearer or farther, so the layer becomes a shallow sculpture instead of a flat sheet.
- Without it
- Each layer would be a flat sticker. The layers would still slide over each other, but each one would look like cardboard.
Steer the view below and switch between three ways of faking depth. One flat image slides as a whole: no depth. Flat cards slide against each other, but each layer is a sticker. Depth mesh is the component. The lines drawn over the picture are the net itself.
Steer the camera through the layers
InteractiveThe number of cells is the net's resolution: the surface can only bend where there is a point. The component uses 160, which is 25,921 points per layer: enough that the bends are invisible, cheap enough to draw four times a frame. The net is built once and shared by all four layers.
One level deeper. The net is called a mesh, its points are vertices, and the small program that moves them is the vertex shader. The renderer is raw WebGL2 on one canvas, with no libraries. Two lines of the shader are the whole depth effect. The camera shift moves every point by how far it is from the focus depth, so near and far slide in opposite directions and the focus stays put. The perspective line spreads near points apart and draws far ones together, which makes the camera feel like it moves into the scene.
void main() { vUV = vec2(aPos.x, 1.0 - aPos.y); float raw = clamp(texture(uDepth, vUV).r, 0.0, 1.0); float placed = pow(clamp(uDlo + raw * (uDhi - uDlo), 0.0, 1.0), uGamma); vRaw = raw; vDepth = placed; float z = placed - uFocus; vec2 p = (aPos * 2.0 - 1.0) * uScale; p += uCam * z; float s = 1.0 + z * uDolly; vPos = p * s + uLayerOffset + uWindowCenter; gl_Position = vec4(vPos, 0.0, 1.0);}// One shared grid; every layer displaces it by its own depth.const verts = new Float32Array((GRID + 1) * (GRID + 1) * 2);for (let y = 0; y <= GRID; y++) for (let x = 0; x <= GRID; x++) { verts[vi++] = x / GRID; verts[vi++] = y / GRID; }// two triangles per cell, Uint32 indices (WebGL2)idx[ii++] = a; idx[ii++] = c; idx[ii++] = b;idx[ii++] = b; idx[ii++] = c; idx[ii++] = d;A window, not a crop
- What it does
- The rounded frame hides everything outside it, except the hand with the glass, which is allowed to cross it.
- Why it exists
- An ordinary crop cuts off everything in the box, with no exceptions. Here the cut is decided layer by layer, so one layer can be let through.
The gold-ringed frame you see is only a drawing of a frame: a dark fill and two thin rings. It does not cut anything. The painting is drawn on a larger canvas that covers the whole stage, and for every pixel of every layer the renderer asks one question: is this pixel inside the frame? The glass layer simply skips the question.
Clip with a distance field
InteractiveOne level deeper. The question is answered with a signed distance: how far the pixel is from the rounded rectangle's edge, negative inside. Along the sides it is a straight distance; near a corner it is the distance to the corner's circle. A soft step across 1.5 pixels turns it into a smooth edge at any size.
float mask = 1.0;if (uWindowClip > 0.5) { vec2 px = (vPos - uWindowCenter) * uResolution * 0.5; vec2 q = abs(px) - uWindowHalf + uWindowRadius; float edge = length(max(q, 0.0)) + min(max(q.x, q.y), 0.0) - uWindowRadius; mask = 1.0 - smoothstep(-0.75, 0.75, edge);}outColor = c * uFade * mask;// per layer, on the CPU side:gl.uniform1f(U.uWindowClip, foreground && params.escape !== false ? 0 : 1);Scroll opens the window
- What it does
- As you scroll, the frame grows wider and taller, the painting zooms in a little, and the hand reaches toward you.
Everything in the opening follows one number, the reveal, which goes from 0 (closed) to 1 (open). It starts when the top of the painting has risen to 70% of the way down the screen and finishes when it reaches 5%. Drag the scroll below and watch the numbers move together.
Scrub the scroll
InteractiveThe window widens from 78% to 86% of the stage and its height grows from 54% to 64% of the stage's width. The whole scene zooms 5.5%. The perspective goes from a quarter of its strength to all of it, so the painting starts flat and deepens as it opens. And the glass scales from 0.88 to 1.10 with a small lift, which turns a growing crop into a gesture: the hand reaches toward you.
One level deeper. The reveal is not a straight line. It is smoothed (a smoothstep) so it starts gently and settles gently instead of snapping on and off at the two lines. And it is measured, not animated: every frame the engine reads where the painting is on the screen and computes the number from that, so scrolling back closes the window again.
// Reveal: from the stage's top at revealStart of the window to revealEnd,// smoothstepped. It drives the window's size, the zoom, and the reach.const stageTop = stageEl.getBoundingClientRect().top - rootRect.top;const rs = num(params.revealStart, 0.7);const re = num(params.revealEnd, 0.05);const lin = clamp01((rs * H - stageTop) / Math.max(1, (rs - re) * H));const reveal = reduce ? 0 : lin * lin * (3 - 2 * lin);windowEl.style.width = `${lerp(wFrom, num(params.widthTo, 86), reveal)}%`;windowEl.style.height = `${lerp(hFrom, num(params.heightTo, 64), reveal)}cqw`;const sceneZoom = 1 + reveal * num(params.sceneZoom, 0.055);// The near hand starts recessed; its zoom and lift turn the reveal// into a reach rather than just a growing crop. It alone escapes.const foreground = l.name === 'drink';const emerge = foreground ? reachFrom + reveal * reach : 1;gl.uniform2f(U.uScale, sx * guard * sceneZoom * emerge, sy * guard * sceneZoom * emerge);gl.uniform2f(U.uLayerOffset, foreground ? -reveal * 0.012 : 0, foreground ? reveal * 0.016 : 0);// and perspective builds with itconst dolNow = reduce ? 0 : dolly * cur.amp * 6 * (0.25 + reveal * 0.75);A candle in front of the paint
- What it does
- A warm light follows the cursor and brushes across the paint, catching on folds of fabric, the glass and the stonework.
- Why it exists
- It makes the painting feel like a physical surface in a room, and it costs no extra images: it reuses the depth maps that are already loaded.
A depth map says how near each point is. Compare a point with its neighbours and you also know which way the surface is facing there: toward the light or away from it. That is all a light needs. The lamp hovers just in front of the scene at the pointer, and each pixel is brightened by how directly it faces the lamp and how close it is.
Carry the candle
InteractiveOne level deeper. The lab opens on light only and lambert: the light pattern alone, with the textbook formula. Carry the lamp to the edge of the glass and the far side cuts to black. Wrapped fixes that hard edge: 0.7 · N·L + 0.3 softens the shadow side instead of switching it off. Switch to lit and the lamp meets the paint: the light multiplies the paint's own colour and is soft-capped, so the paint is warmed and lifted, never fogged. The falloff is gentle, 1 / (1 + 1.5·d²), so there is no hot spot. And there are three lamps, not one: tight and gentle on the glass (gain 0.85), strong on the figure (1.6), broad and faint on the background (0.6).
vec3 n = normalize(vec3((dl - dr) * uRelief, (dd - du) * uRelief, 1.0));float fragZ = (1.0 - vDepth) * 0.9;vec3 Lv = vec3(uLight + uLightOff - vPos, 0.25 + fragZ);float d2 = dot(Lv.xy, Lv.xy) * uLightK + Lv.z * Lv.z * 0.35;float atten = 1.0 / (1.0 + d2 * 1.5); // fat tail, gentle peakvec3 Ld = Lv / max(length(Lv), 1e-4);float diff = clamp(dot(n, Ld) * 0.7 + 0.3, 0.0, 1.0); // wrapped: no hard terminatorfloat I = atten * uLightAmp * uLightGain;c.rgb *= 1.0 - uRoomDim * clamp(uLightAmp, 0.0, 1.0);float lift = diff * I * uLightDiff * guard;lift = lift / (1.0 + 0.6 * lift);c.rgb += c.rgb * uLightColor * lift; // multiplicative: warmed, never foggedvec3 H = normalize(Ld + vec3(0.0, 0.0, 1.0));float spec = pow(max(dot(n, H), 0.0), 20.0);c.rgb += uLightColor * spec * I * uLightSpec * guard * c.a;c.rgb = min(c.rgb, vec3(c.a));The normals view is the one to remember: it shows what the light actually has to work with, and that texture is the depth map itself. The glass and the drapery are rich because their relief bends; the face is nearly flat because its map is.
Depth from scrolling alone
- What it does
- The view drifts slowly as the section scrolls past, and the depth effect is strongest while the painting is in the middle of the screen.
- Why it exists
- Depth only shows when the view moves, and otherwise the view only moves with the cursor. A phone has no cursor, and many desktop visitors never move theirs over the painting. Scrolling is the one thing everybody does.
This is separate from the window opening. The opening changes the frame; this moves the camera. It is two small effects, both driven by how far the section has travelled through the screen (0 as it enters at the bottom, 1 as it leaves at the top):
- The drift. The camera moves vertically with the scroll, so near layers slide one way and far ones the other as the page goes by, like walking past a window. This is the part that works without a cursor.
- The swell. The camera's range is
0.075 + 0.05 · sin(progress · π): smallest as the section arrives and leaves, two thirds larger in the middle. The painting is calm at the edges of the screen and most alive when it is centred.
What section progress does
InteractiveBoth are deliberately small. At the component's numbers, on a stage about 830 px tall, the glass moves about 20 px against the sky over the whole scroll: you sense it as depth rather than see it as motion. That is why the figure opens at ×4.
// Depth: the section's whole progress, 0 as its top enters the bottom of// the window, 1 as its bottom leaves the top. Parallax blooms mid-way.const sp = clamp01((H - (secRect.top - rootRect.top)) / (secRect.height + H));tgt.amp = baseAmp + boost * Math.sin(sp * Math.PI);tgt.tilt = (sp - 0.5) * tiltAmt;// …const camY = reduce ? 0 : cur.y * cur.amp * 0.6 + cur.tilt * cur.amp;Following, not snapping
- What it does
- Nothing jumps to where the cursor is. The camera and the light glide there, and they glide at the same pace on every screen.
- Without it
- Twitchy motion, and a painting that feels heavy on an old laptop and nervous on a fast gaming monitor.
The camera, the swell, the drift and the candle each have a target (where they should be) and a current value (where they are). Every frame the current value closes a small fraction of the gap. The pointer only sets targets; the motion you see is the chase.
Same damping on every display
InteractiveOne level deeper. If every frame closes the same fraction, a screen that draws more frames per second chases faster. With a fraction of 0.065, after one second a 30 Hz screen still has 13% of the distance left, a 60 Hz screen 1.8%, and a 144 Hz screen almost nothing. The engine fixes it by treating 0.065 as the fraction for 1/60 of a second and scaling it to each frame's real length, 1 − (1 − k)^(dt·60), so every display ends the second at the same 1.8%.
/** Frame-rate independent version of "move k of the way each 60fps frame". */function ease(k: number, dt: number): number { return 1 - Math.pow(1 - clamp01(k), dt * 60);}const overArt = pointer.on && Math.abs(pointer.nx) <= 1 && Math.abs(pointer.ny) <= 1;tgt.x = overArt ? -pointer.nx * pStrength : 0;tgt.y = overArt ? pointer.ny * pStrength * 0.75 : 0;const k = ease(num(params.damping, 0.065), dt);cur.x += (tgt.x - cur.x) * k;cur.y += (tgt.y - cur.y) * k;Soft edges without fringes
- What it does
- Where a cut-out's edge is soft (hair, sheer fabric, fingertips), the component mixes it into the layer behind in a way that keeps the edge clean.
- Without it
- A faint halo around the model's hair and fingers, most visible where they cross the sky.
The figure blends the real layers three ways, zoomed in so you can see single pixels. The first option is what the component does; the other two are the two ways of getting it wrong.
Why the colour is premultiplied
InteractiveOne level deeper. How see-through a pixel is, is called its alpha. The component multiplies each colour by its own alpha once, as the image is decoded (premultiplied alpha), then blends with ONE, ONE_MINUS_SRC_ALPHA. Get either half wrong and it shows: skip the multiply and soft edges glow bright; multiply twice and they go dark. Depth maps are never premultiplied, because they are measurements, not colours.
// colour is premultiplied as it is decoded, depth is left raw:// premultiplying a depth map would darken its relief at soft edgesloadBitmap(ASSETS + l.color, true), loadBitmap(ASSETS + l.depth, false)async function loadBitmap(src: string, premultiply: boolean): Promise<ImageBitmap> { const response = await fetch(src); return createImageBitmap(await response.blob(), { premultiplyAlpha: premultiply ? 'premultiply' : 'none', colorSpaceConversion: 'none', });}gl.enable(gl.BLEND);gl.blendFunc(gl.ONE, gl.ONE_MINUS_SRC_ALPHA);Never show the edge
- What it does
- The sky is drawn slightly bigger than the frame: about 14% bigger on the default settings.
- Why it exists
- When the view moves, the sky slides sideways and shrinks a little. If it were exactly the size of the frame, its edge would come into view.
The sky is the back wall of the scene, and behind it there is nothing. In the figure the camera is pushed as far as the component ever lets it go, and the building is switched off so you can see the sky on its own. The dashed line is the sky's real edge, and the striped area is empty frame showing through.
Keep the sky behind the frame
InteractiveNo guard. The sky is exactly as big as the frame, so any camera move pulls its edge inside.
Now switch on show the building. The arch covers almost the whole border of the frame, so in this painting the gap is nearly always hidden anyway. The guard is a safety net: it is what keeps the frame full when the building layer is hidden, when the depth is turned up, or when the same component is given different artwork.
One level deeper. Two things pull the sky's edge inward. The camera shift slides it, by up to 4.5% of the stage's half width. And the perspective shrinks it, by about 8%, because it is far away and far things are drawn smaller. The guard undoes both: (1 + 0.045) / 0.919 = 1.137. It is not a fixed number. The engine recomputes it every frame from the current settings, so it stays exact when someone turns the depth up.
Why only the sky? Because scaling a layer also moves everything painted on it. Switch the figure to guard every layer: the glass and the figure grow and drift out of the composition. The arch does not need the guard either, since it sits in front of the focus, where perspective makes things bigger, not smaller.
// Worst-case motion bounds, for the sky's edge guard below.const maxAmp = baseAmp + boost;const dolMax = dolly * maxAmp * 6;const camMax = maxAmp * (pStrength + tiltAmt * 0.5);// Edge guard: only the sky is scaled up by what its worst-case motion// could expose; enlarging the cutouts would crop the cup.const zAbs = Math.max(Math.abs(lo - focus), Math.abs(hi - focus));const sMin = Math.max(0.6, 1 - zAbs * dolMax);const guard = l.name === 'sky' ? (1 + camMax * zAbs) / sMin : 1;Letters that rise
- What it does
- The heading's letters rise and fade in one after another when the heading scrolls into view, and again whenever you scroll back above it and return.
The heading waits until its fonts have loaded and its top crosses 72% of the screen, then plays. Three choices shape it, and each is a control in the figure:
- Motion ease: how one letter travels to rest. Expo out shoots up and settles; back out overshoots; linear is mechanical.
- Cascade curve: when each letter starts. Even spaces the starts equally; speeds up opens with wide gaps that close as it goes; slows down does the opposite.
- First and last letter durations: the same curve also moves each letter's duration from the first letter's to the last letter's, so a cascade can start slow and finish fast.
Play the heading, one letter at a time
InteractiveHeavenhasamenu
One bar per letter, from the moment it starts to the moment it lands. Where the bars begin is the cascade; how long they are is each letter's duration. The whole heading takes 1.46 s.
// heading-timing.ts, shared by the engine and the lab aboveexport function letterTimings(n: number, options: HeadingTimingOptions): LetterTiming[] { const curve = HEADING_CASCADES[options.cascade] ?? HEADING_CASCADES.even; const spread = Math.max(0, options.stagger) * Math.max(0, n - 1); return Array.from({ length: n }, (_, i) => { const t = n <= 1 ? 0 : i / (n - 1); const c = curve(t); // one curve for both: when it starts, how long it takes return { delay: Math.round(spread * c), duration: Math.round(options.durationFirst + (options.durationLast - options.durationFirst) * c), }; });}// engine.ts, playReveal()chars.forEach((ch, i) => { ch.animate( [{ transform: `translateY(${rise}%)`, opacity: 0 }, { transform: 'translateY(0)', opacity: 1 }], { duration: timings[i].duration, delay: timings[i].delay, easing, fill: 'backwards' }, );});One level deeper. To animate letters separately the heading is split into one box per letter, and separate boxes lose the font's fine spacing between letter pairs. So when the last letter lands, the heading is turned back into plain text. Scroll back up past the 72% line and it fades out and is split again, ready to play the next time it comes up. All of this is in the component's Section folder, and Replay heading plays it again after an edit.
Sized by the screen
- What it does
- Every size in the section is a share of the screen's height: the empty space around it, the painting's height, and the points where the opening starts and ends.
- Without it
- On a short laptop the painting is taller than the screen and starts opening before you can see it. On a tall monitor it is small and opens late.
The figure shows a web page on four screens. With a fixed 900 px, the section keeps the sizes it was designed with on one desktop. With this screen, it measures the screen it is actually on. Switch between them on the short laptop first.
One section, four screens
InteractiveThe rules are simple: half a screen of empty space above and below, a painting never taller than 92% of the screen, an opening that runs from 70% to 5%, a heading that plays at 72%.
One level deeper. CSS already has a unit for this, vh, but it always means the browser window. This section does not assume it owns the whole window. It measures the height of whatever box it scrolls in and writes it into a CSS variable, --dw-vh, and every size uses that. On a normal page the box is the browser window and the result is identical. Inside a scrolling panel, a modal, or the small preview at the top of this article, it is the difference between fitting and overflowing. For the same reason the layout switches at a section width of 768 px with a container query, not a media query.
.dw-space { height: calc(var(--dw-vh, 800px) * var(--dw-space, 0.5)); }.dw-stage { container-type: inline-size; width: min(100%, calc(var(--dw-vh, 800px) * 0.92 / 0.82)); height: min(82cqw, calc(var(--dw-vh, 800px) * 0.92));}@container dw (min-width: 768px) { .dw-section { padding: 144px 5% 112px; }}// every frameconst h = host.clientHeight;if (h && h !== viewH) root.style.setProperty('--dw-vh', `${h}px`);Two fallbacks complete it. With reduced motion switched on, the window stays closed, the camera holds still, the candle is out and the heading appears as plain text. Without WebGL2, the canvas is replaced by a flat image of the painting.
How a model sees depth
- What it does
- The depth maps were not painted by hand. A machine-learning model, Apple's Depth Pro, looked at each layer and estimated how far away every pixel is.
- Why it exists
- Measuring depth normally takes two cameras or a laser scanner. A painting has neither: there is only one flat picture, so its depth has to be estimated from what is in it.
How can a program judge distance from a single picture? Nobody wrote it rules such as "small means far". It was trained: shown a very large number of images for which the true distance of every pixel was already known, some of them real scenes with measured depth and some computer-generated scenes where the depth is known exactly. For each image the model makes its estimate, the estimate is compared with the truth pixel by pixel, and the model's internal numbers are nudged so the error gets a little smaller. Repeated enough times, it becomes good at pictures it has never seen, including this painting.
What exactly it looks for is not written down anywhere, and the paper does not list it. Presumably it is the same kind of evidence you use with one eye closed: things that overlap, things that shrink with distance, the way light falls. What is known is how it looks at the picture, how its answer becomes gray, and how it was used here. Those are the three steps below.
Step 1: three looks at the same picture. The model resizes the picture to a square of 1536 pixels, then makes two smaller copies, 768 and 384 pixels across. It cuts all three into squares of 384 pixels: 25 from the large copy, 9 from the middle one, and the small copy is a single square. That is 35 patches, and they overlap so no seams appear. The small copy shows the whole scene, blurred; the large copy shows every detail, but only a sixteenth of the picture at a time.
Three looks at the same picture
InteractiveDetail. Twenty-five overlapping patches, each a close-up. Hair and fingers are sharp, but a single patch has little idea where it sits in the scene.
The same reader goes over every patch. It divides each one into a grid of 24 by 24 small squares and lets every square compare itself with all the others in the patch, so the answer for a strand of hair can depend on the shoulder next to it. The readings from all 35 patches are then merged back into one map, so the result can draw on both the overview of the whole scene and the detail of the close-ups. The result is a depth value for every pixel of the 1536 pixel square, which the paper reports takes about 0.3 seconds on a standard graphics card.
Step 2: from distance to gray. The model's answer is not stored as plain distance. It is stored as one divided by the distance, which is large for near things and close to zero for far ones. The pipeline that made these maps then stretches the result so the farthest point becomes black (0) and the nearest becomes white (1). The figure shows why, using example distances.
From a distance to a shade of gray
InteractiveWhat the model and the pipeline use. Half of all the gray is spent between 0.5 m and 1 m. Near things get the detail; everything far goes dark together.
This is why the maps look the way they do: bright, detailed shapes close to you, and backgrounds that fade into the same dark gray. It suits this component, because near things are the ones that move the most.
Step 3: one pass per layer. The model expects a picture of a scene, not a cut-out floating on nothing; it could read the emptiness around a cut-out as endlessly far away. So each layer was shown to it sitting on top of the layers behind it: first the sky alone, then sky and building, then those two with the woman, then all four. From each answer only the part under that layer's own outline was kept.
One pass per layer
InteractiveLeft: what the model was shown. Right: the part of its answer that was kept, cut to the figure's outline and stretched so its farthest point is black and its nearest is white.
One level deeper. The reader is a Vision Transformer, and the comparing step is called attention. Running attention over the full picture at once would be far too costly, which is why the model works in 384 pixel patches. It learns in two stages: first on a mix of real and computer-generated images, to cope with any kind of picture; then on computer-generated images only, because real measured depth has gaps and errors exactly along object edges. In that second stage it is also graded on how well its edges line up, which is why hair and fingers come out crisp instead of smeared.
The model can also estimate the camera's focal length, which turns its answer into real metres: distance = focal length / (width · C), where C is the model's output. That part is not needed here. The pipeline keeps only the order and proportions: it takes one over the distance, stretches it to 0 to 1, snaps the result to the picture's own edges with a guided filter, cuts it to the layer's outline, and stretches it again over that outline alone, so each map describes a shape and nothing about where the shape sits. Placing it is the job of section 4.
Sources
The local projects the component was rebuilt from, and the published work they stood on. Local paths are relative to the Relume-Content workspace.
The projects it was rebuilt from
- Matcha Angels (Design 2, “Holy Green”): the Shrine sectionRelume-Content/Test-build-1-relume · src/pages/Design2.jsx (Shrine), src/index.css (.shrine-*)The section this component copies: the eyebrow, heading, caption and copy; the stage sizing (min(82cqw, 92vh)), the window’s 4cqw offset, gold hairline rings and radius; the widening window, originally driven by Motion’s useScroll from ‘start 85%’ to ‘start 5%’ (here the start is 70%).
- DepthHero.jsx, the Shrine’s rendererRelume-Content/Test-build-1-relume · src/components/DepthHero.jsxThe WebGL2 core, nearly line for line: the 160-segment grid, band placement, painter’s-order premultiplied blend, the rounded-window SDF clip with the drink exempt, the reveal’s zoom, dolly ramp and reach, the bloom and tilt, the sky-only edge guard, and the three-group candle lamp with its constants.
- DepthDials.jsx and textReveal.jsRelume-Content/Test-build-1-relume · src/components/DepthDials.jsx, src/lib/textReveal.jsThe Depth and Lamp folders of the control panel and their ranges; the heading reveal’s timing: 42% rise, 24 ms stagger, 1.15 s, expo out, once at ‘top 72%’, reverted after. Originally GSAP SplitText, rebuilt here on the Web Animations API.
- depth3d-v7, per-layer live controlRelume-Content/matcha-renaissance/depth3d-v7 · src/Controls.jsx, README.md, spec.jsonThe per-layer folders (visible, lo, hi, show depth), the band seeds (sky 0–0.10, building 0.60–0.85, model 0.20–0.55, drink 0.87–1.00), the draw order correction, and the gamma, focus, grain, vignette and aberration dials.
- depth3d-v4, the layered depth rendererRelume-Content/matcha-renaissance/depth3d-v4 · depth-hero4.jsThe lineage of the renderer: mesh-mode displacement, far-anchored parallax, a positional lamp rather than a directional one, premultiplied alpha end to end, painter’s order with the depth test off.
- PLANV6.md, relief.py, build.py: the asset pipelineRelume-Content/matcha-renaissance · PLANV6.md, depth3d-v7/relief.py, depth3d-v7/build.pyWhy each layer gets its own depth pass (Change A), why shape and placement are separate (Change B), why the depth is flooded past each silhouette (Change C); Depth Pro at its 1536 px working size, output at 2200 px, lossless WebP with exact colour under transparency.
Published references
- Depth Pro: Sharp Monocular Metric Depth in Less Than a Second (Bochkovskiy et al., Apple, 2024)arxiv.org/abs/2410.02073 · github.com/apple/ml-depth-pro · huggingface.co/apple/DepthPro-hfThe depth model every layer’s map was produced with (run through the Hugging Face port in relief.py), and the source for how it works: the fixed 1536 × 1536 input, the 35 overlapping 384 px patches at three scales, the 24 × 24 features per patch, inverse depth as the output, the focal-length formula, the two training stages, and the 0.3 s runtime.
- Depth map, Wikipediaen.wikipedia.org/wiki/Depth_mapThe definition the section opens with: an image or image channel that records the distance of a scene’s surfaces from a viewpoint. Here every layer carries its own, brighter meaning nearer.
- Heightmap, Wikipediaen.wikipedia.org/wiki/HeightmapThe idea the depth-map lab shows: a grayscale image read as height, so one value per pixel lifts a flat grid into a surface.
- Displacement mapping, Wikipediaen.wikipedia.org/wiki/Displacement_mappingWhy relief can look soft up close: displacement moves real geometry, and its quality is limited by the density of the mesh, so the surface bends only where the grid has a vertex.
- Premultiplied alpha, Shawn Hargreavesshawnhargreaves.com/blog/premultiplied-alpha.htmlWhy straight alpha fringes at soft edges and premultiplied alpha doesn't, the failure the fringe lab reproduces.
- 2D distance functions, Inigo Quileziquilezles.org/articles/distfunctions2dThe rounded-box signed distance the window clip evaluates per fragment.
- Alpha compositing and premultiplied alphaen.wikipedia.org/wiki/Alpha_compositing · MDN, createImageBitmapThe “over” operator the blend implements. The premultiply itself happens in createImageBitmap, as each colour layer is decoded.
- WebGL2RenderingContext, MDNdeveloper.mozilla.org/en-US/docs/Web/API/WebGL2RenderingContextThe context the engine opens with one getContext('webgl2') call on a plain canvas: OpenGL ES 3.0 shaders, vertex arrays, Uint32 indices, and no rendering library anywhere in the component.
- Half Lambert, Valve Developer Communitydeveloper.valvesoftware.com/wiki/Half_LambertThe idea behind wrapped diffuse: remap N·L so the terminator softens instead of cutting to black (this lamp uses 0.7·N·L + 0.3).
- Blinn–Phong reflection modelen.wikipedia.org/wiki/Blinn–Phong_reflection_modelThe half-vector specular glint: normalize(L + V), raised to the 20th power.
- Frame rate independent damping using lerp, Rory Driscollrorydriscoll.com/2016/03/07/frame-rate-independent-damping-using-lerpWhy cur += (tgt − cur)·k drifts with frame rate, and the exponential correction the ease() helper applies.
- Element.animate(), MDN, Web Animations APIdeveloper.mozilla.org/en-US/docs/Web/API/Element/animateThe dependency-free replacement for the original GSAP tween: per-letter keyframes, delay for the stagger, fill backwards.
- SplitText (GSAP docs) and useScroll (Motion docs)gsap.com/docs/v3/Plugins/SplitText · motion.dev/docs/react-use-scrollWhat the original section used for the heading split and the scroll-linked window; this component keeps their behaviour and drops both dependencies.
- easeOutExpo, easings.neteasings.net/#easeOutExpocubic-bezier(0.16, 1, 0.3, 1), the curve the letters rise on.
- Container queries, MDNdeveloper.mozilla.org/en-US/docs/Web/CSS/CSS_containment/Container_queriescqw units for the stage and window, and the @container rule that stands in for the original’s md: breakpoint.
- perspective and transform-style, MDNdeveloper.mozilla.org/en-US/docs/Web/CSS/perspective · developer.mozilla.org/en-US/docs/Web/CSS/transform-styleThe CSS 3D the peel lab is built on: one perspective on the stage, preserve-3d on the stack, rotateY and translateZ per plane.
- DialKit, Josh Puckettgithub.com/joshpuckett/dialkitThe control panel, in depth3d-v7 and here: every number in this article is a dial on the component page.
- Instrument Serif, Instrument Sans, EB Garamond (Google Fonts)fonts.google.com/specimen/Instrument+SerifThe section’s three faces: display heading, eyebrow, italic caption.
Make it yours
That is the whole machine: four layers and four depth maps; a band that gives each shape its place; a net of points bent by depth; clean soft edges; a frame that is a rule one layer may ignore; an opening, a drift and a guard measured from the scroll every frame; a camera that glides; a candle that warms the paint; a heading that rises as it enters; and sizes that follow the screen.
Every number in this article is a dial on the component's page, with its build prompt beside it. Start with the presets: Candlelit for a warmer room, Deep parallax for more movement, Big reach for a hand that comes further out, Gallery to keep everything inside the frame, and Depth view to see the four depth maps doing the work.