github tB0nE/nightfall edgepad-models-2026-10-08
EdgePad depth models (ZipDepth for mobile GPUs)

pre-release6 hours ago

The monocular depth models Nightfall uses for AI 3D, published for anyone to use: TFLite for phones and headsets (Nightfall runs them on Quest), ONNX for PCs (Nightfall Meteor runs them there). MIT licensed, see LICENSE.

All of them are EdgePad builds of ZipDepth by Fabio Tosi (MIT): a 6.1M-parameter convolutional network distilled from Depth Anything V2 Large, with no attention or transformer ops, which is what lets it run entirely on a mobile GPU. There are two generations:

  • Square EdgePad (shipped in Nightfall v0.7.x): ZipDepth's own Standard checkpoint, not retrained, exported through exact graph rewrites so the whole network runs on the Quest's GPU.
  • Widescreen EdgePad (Nightfall's next release): a light fine-tune of the same checkpoint for 16:9 video and steadier depth from frame to frame, exported at exact 16:9 sizes.

Files

File Generation Input Output Used by SHA-256
zipdepth-base-384-standard-packed-conv4-reduceconv-edgepad-gpu.tflite Square 384x384 packed 192x192x4 Quest 3 default (v0.7.x) 6a2f5b073f2650768c74921156dad07c7c61006920f41f89baa52e076a3aef25
zipdepth-base-256-standard-packed-conv4-reduceconv-edgepad-gpu.tflite Square 256x256 packed 128x128x4 Quest 2 default (v0.7.x) 8a636c3551e35886172811e41a03a2b6e5c8e3398037a0a0b58ca25a3332076f
zipdepth-base-224-standard-packed-conv4-reduceconv-edgepad-gpu.tflite Square 224x224 packed 112x112x4 Manual option (v0.7.x) 8d88be4a2ffe0be29539cb0d1364ee3a83ec82f3a0a2d5d7cc25f157752fb9ff
zipdepth-base-384-cpu.tflite Square 384x384 384x384x1 CPU backend 4db1d7ed7d0f6883124be97b7ef02e5ebed2e442b4d59240eeeaedf8135169c6
zipdepth-base-256-cpu.tflite Square 256x256 256x256x1 CPU backend bcca2459c6e97b7bd15c009175fd7372b320522d70627f958590a4385576cbf4
zipdepth_edgepad_384.onnx Square 384x384 384x384 Meteor (PC) 78efcf08caded436c626af0b7b5980db47bc07653ec549a59b7e48e73b3e035f
zipdepth-wide-512x288-edgepad-gpu.tflite Widescreen 512x288 packed 256x144x4 Quest 3 default (next release) dfa40c15155df2a7ae2eb2ccdf78da610ae4722b1cc388fa93e6888bcc78d0df
zipdepth-wide-320x180-t320x192-edgepad-gpu.tflite Widescreen 320x180 in a 320x192 tensor packed 160x96x4 Quest 2 default (next release) 9dfd7519b3d1b208b7ad74caa55b9097e35ca537b03fdd0ec690d8cfce56f2ec
zipdepth_wide_512x288.onnx Widescreen 512x288 512x288 Meteor default (PC) 197d598f50ee3270b12831096dbab260c6bd5979bab13b87f01a28fad0268274
zipdepth_wide_672x384.onnx Widescreen 672x384 672x384 Meteor, larger option (PC) f812b10b10309c6759266cbe888fa9c0f70668350fdf69e7eda9cc0412ca1a21

Sizes are width x height. The packed outputs are listed as width x height x channels.

Using them

  • Input: RGB as float32 in 0 to 1 (pixel / 255). ImageNet mean and standard deviation are applied inside the graph. TFLite takes NHWC [1, H, W, 3]; ONNX takes NCHW [1, 3, H, W] named image. Resize the frame to the input size (Nightfall uses an area filter).
  • Output: relative inverse depth: larger is nearer, with no fixed scale or offset. Normalise each frame before use (Nightfall stretches between low and high percentiles, smoothed over time).
  • Packed outputs (the GPU TFLite models): each output pixel's 4 channels are a 2x2 block of the full-size map, in the order top-left, top-right, bottom-left, bottom-right. Interleave them to twice the width and height, then clamp at zero (ReLU). This replaces the original graph's softmax and depth-to-space tail, which split the graph off the GPU. The CPU and ONNX models already output the full-size map, unpacked and clamped.
  • 320x180 in 320x192: place the 320x180 image in the middle of the tensor and fill the 6 rows above and below by repeating its first and last rows. Unpack the output to 320x192, then crop rows 6 to 185. Cropping the packed tensor instead misaligns the 2x2 blocks. Stretching the image to 320x192 is wrong.
  • Precision: the GPU TFLite models store float16 weights with float32 input and output. Nightfall runs them on LiteRT's GPU delegate (OpenCL on Quest 3, OpenGL on Quest 2) with precision loss allowed. The CPU models are int8 weights with float32 activations (w8a32), for XNNPACK with no GPU at all. The ONNX models are float32.

Square EdgePad: optimisation only

The weights are ZipDepth's Standard checkpoint (zipdepth_base.pth), unchanged, including its learned convex reconstruction head, which proved essential for clean full-screen geometry in the headset. Every change below is to the graph and is mathematically exact:

  1. Materialised attention broadcasts. ZipDepth's strip, channel and global-context attention add or multiply small tensors that broadcast across the feature map. That's legal TFLite and passes the delegate's compatibility check, but the Quest 3's Adreno OpenCL backend computes it wrongly. The small tensors are expanded to full size (nearest-neighbour resize) before the operation.
  2. Small reductions. Large spatial average pools are split into exact 2x and 3x stages.
  3. Packed reconstruction head. The head's 5-dimensional softmax and depth-to-space tail becomes four ordinary 9-channel softmax branches, the 36-channel mask convolution becomes four 9-channel convolutions, and the weighted sums become fixed 1x1 convolutions. The output is the packed 2x2 layout described above.
  4. EdgePad. The head's replicate padding (MIRROR_PAD) isn't supported by the GPU delegate. It split the graph and copied about 5 MiB of intermediate tensors back to the CPU on every inference. Swapping in zero padding proved the cost: the whole graph ran in one OpenCL partition and a 29 ms build dropped to about 18 ms on Quest 3, but the outermost pixels changed. EdgePad rebuilds the exact replicate padding from edge slices and concatenation, which the delegate runs, keeping both the speed and the exact output.

Against the original Standard graph: correlation 0.9999999999999908, normalised mean absolute error 3.6e-8. Steady inference on Quest 3 with OpenCL: about 18 ms at 384x384 and 15 ms at 256x256. The CPU models quantise the same 384 and 256 exports to w8a32.

Widescreen EdgePad: a light fine-tune

One checkpoint (SHA-256 2dc8fddcef9ef22f45b6a1573b9cf66de0f39b1b5172e4284dd693790bd6309e), trained from ZipDepth's Standard checkpoint in two short stages. Only the decoder was trained: 154,005 of 6.1 million parameters. The encoder is ZipDepth's.

  • Stage 1 (5,000 steps, 170 s on an RTX 3090): shapes alternated between 384x384, 256x256, 224x224 and 512x288, so one set of weights serves every size.
    • Data: 2,000 frames from 25 videos: 10 gameplay and YouTube, 8 film, live-action or CG, 7 animation. The frames came as 250 clips of 8 consecutive frames, with scene cuts marked so no temporal loss crosses one.
    • Labels: the original square EdgePad-384's own output on every frame (an anchor, so the model doesn't drift from ZipDepth); Video Depth Anything Small (Apache-2.0) for frame-to-frame consistency on every clip; and Marigold V2 (Apache-2.0) on one hard keyframe per clip for spatial detail.
    • A texture-leakage term at weight 0.5 discourages copying image texture into the depth.
    • Temporal behaviour is learned only through the losses: the model has no memory, and each frame is still processed on its own.
  • Stage 2 (1,000 steps at learning rate 5e-5, 44 s): decoder only again, for the exact 16:9 sizes.

The training frames aren't distributed.

Results, on held-out video against the untouched ZipDepth export at the same size (negative is better):

Model Compared with Flicker on static scenes Flow-warped error Frame-to-frame change
512x288 stock 512x288 -8.6% -5.8% -8.4%
320x180 stock 224x224 -42.4% -43.9% -39.2%

The cost is texture leakage: about 11% higher than the stock 224x224 at the 320x180 size. The 320x180 model is also visibly less detailed than the 384x384 one, as its size implies.

Exports: the TFLite models match their ONNX source with correlation 0.99999 (512x288 and 320x192), and LiteRT's GPU compatibility check reports no warnings. Meteor's ONNX files are the same weights with the 2x2 unpacking and clamp added to the graph and the weights embedded. The 672x384 file is an export of the same checkpoint at a larger size; it wasn't part of the held-out evaluation above.

Credits and licences

  • ZipDepth: Fabio Tosi, MIT. These models are derivatives of its weights; its copyright notice is kept in LICENSE.
  • Teacher models used only to make training labels, not included: Video Depth Anything Small (Apache-2.0) and Marigold V2 (Apache-2.0; built on Qwen-Image-Edit-2509, Apache-2.0).
  • Fine-tuning, graph optimisation and exports: Nightfall's model research.

Don't miss a new nightfall release

NewReleases is sending notifications on new releases.