The monocular depth models Nightfall uses for AI 3D, published for anyone to use: TFLite for phones and headsets (Nightfall runs them on Quest), ONNX for PCs (Nightfall Meteor runs them there). MIT licensed, see LICENSE.
All of them are EdgePad builds of ZipDepth by Fabio Tosi (MIT): a 6.1M-parameter convolutional network distilled from Depth Anything V2 Large, with no attention or transformer ops, which is what lets it run entirely on a mobile GPU. There are two generations:
- Square EdgePad (shipped in Nightfall v0.7.x): ZipDepth's own Standard checkpoint, not retrained, exported through exact graph rewrites so the whole network runs on the Quest's GPU.
- Widescreen EdgePad (Nightfall's next release): a light fine-tune of the same checkpoint for 16:9 video and steadier depth from frame to frame, exported at exact 16:9 sizes.
Files
| File | Generation | Input | Output | Used by | SHA-256 |
|---|---|---|---|---|---|
zipdepth-base-384-standard-packed-conv4-reduceconv-edgepad-gpu.tflite
| Square | 384x384 | packed 192x192x4 | Quest 3 default (v0.7.x) | 6a2f5b073f2650768c74921156dad07c7c61006920f41f89baa52e076a3aef25
|
zipdepth-base-256-standard-packed-conv4-reduceconv-edgepad-gpu.tflite
| Square | 256x256 | packed 128x128x4 | Quest 2 default (v0.7.x) | 8a636c3551e35886172811e41a03a2b6e5c8e3398037a0a0b58ca25a3332076f
|
zipdepth-base-224-standard-packed-conv4-reduceconv-edgepad-gpu.tflite
| Square | 224x224 | packed 112x112x4 | Manual option (v0.7.x) | 8d88be4a2ffe0be29539cb0d1364ee3a83ec82f3a0a2d5d7cc25f157752fb9ff
|
zipdepth-base-384-cpu.tflite
| Square | 384x384 | 384x384x1 | CPU backend | 4db1d7ed7d0f6883124be97b7ef02e5ebed2e442b4d59240eeeaedf8135169c6
|
zipdepth-base-256-cpu.tflite
| Square | 256x256 | 256x256x1 | CPU backend | bcca2459c6e97b7bd15c009175fd7372b320522d70627f958590a4385576cbf4
|
zipdepth_edgepad_384.onnx
| Square | 384x384 | 384x384 | Meteor (PC) | 78efcf08caded436c626af0b7b5980db47bc07653ec549a59b7e48e73b3e035f
|
zipdepth-wide-512x288-edgepad-gpu.tflite
| Widescreen | 512x288 | packed 256x144x4 | Quest 3 default (next release) | dfa40c15155df2a7ae2eb2ccdf78da610ae4722b1cc388fa93e6888bcc78d0df
|
zipdepth-wide-320x180-t320x192-edgepad-gpu.tflite
| Widescreen | 320x180 in a 320x192 tensor | packed 160x96x4 | Quest 2 default (next release) | 9dfd7519b3d1b208b7ad74caa55b9097e35ca537b03fdd0ec690d8cfce56f2ec
|
zipdepth_wide_512x288.onnx
| Widescreen | 512x288 | 512x288 | Meteor default (PC) | 197d598f50ee3270b12831096dbab260c6bd5979bab13b87f01a28fad0268274
|
zipdepth_wide_672x384.onnx
| Widescreen | 672x384 | 672x384 | Meteor, larger option (PC) | f812b10b10309c6759266cbe888fa9c0f70668350fdf69e7eda9cc0412ca1a21
|
Sizes are width x height. The packed outputs are listed as width x height x channels.
Using them
- Input: RGB as float32 in 0 to 1 (pixel / 255). ImageNet mean and standard deviation are applied inside the graph. TFLite takes NHWC
[1, H, W, 3]; ONNX takes NCHW[1, 3, H, W]namedimage. Resize the frame to the input size (Nightfall uses an area filter). - Output: relative inverse depth: larger is nearer, with no fixed scale or offset. Normalise each frame before use (Nightfall stretches between low and high percentiles, smoothed over time).
- Packed outputs (the GPU TFLite models): each output pixel's 4 channels are a 2x2 block of the full-size map, in the order top-left, top-right, bottom-left, bottom-right. Interleave them to twice the width and height, then clamp at zero (ReLU). This replaces the original graph's softmax and depth-to-space tail, which split the graph off the GPU. The CPU and ONNX models already output the full-size map, unpacked and clamped.
- 320x180 in 320x192: place the 320x180 image in the middle of the tensor and fill the 6 rows above and below by repeating its first and last rows. Unpack the output to 320x192, then crop rows 6 to 185. Cropping the packed tensor instead misaligns the 2x2 blocks. Stretching the image to 320x192 is wrong.
- Precision: the GPU TFLite models store float16 weights with float32 input and output. Nightfall runs them on LiteRT's GPU delegate (OpenCL on Quest 3, OpenGL on Quest 2) with precision loss allowed. The CPU models are int8 weights with float32 activations (w8a32), for XNNPACK with no GPU at all. The ONNX models are float32.
Square EdgePad: optimisation only
The weights are ZipDepth's Standard checkpoint (zipdepth_base.pth), unchanged, including its learned convex reconstruction head, which proved essential for clean full-screen geometry in the headset. Every change below is to the graph and is mathematically exact:
- Materialised attention broadcasts. ZipDepth's strip, channel and global-context attention add or multiply small tensors that broadcast across the feature map. That's legal TFLite and passes the delegate's compatibility check, but the Quest 3's Adreno OpenCL backend computes it wrongly. The small tensors are expanded to full size (nearest-neighbour resize) before the operation.
- Small reductions. Large spatial average pools are split into exact 2x and 3x stages.
- Packed reconstruction head. The head's 5-dimensional softmax and depth-to-space tail becomes four ordinary 9-channel softmax branches, the 36-channel mask convolution becomes four 9-channel convolutions, and the weighted sums become fixed 1x1 convolutions. The output is the packed 2x2 layout described above.
- EdgePad. The head's replicate padding (
MIRROR_PAD) isn't supported by the GPU delegate. It split the graph and copied about 5 MiB of intermediate tensors back to the CPU on every inference. Swapping in zero padding proved the cost: the whole graph ran in one OpenCL partition and a 29 ms build dropped to about 18 ms on Quest 3, but the outermost pixels changed. EdgePad rebuilds the exact replicate padding from edge slices and concatenation, which the delegate runs, keeping both the speed and the exact output.
Against the original Standard graph: correlation 0.9999999999999908, normalised mean absolute error 3.6e-8. Steady inference on Quest 3 with OpenCL: about 18 ms at 384x384 and 15 ms at 256x256. The CPU models quantise the same 384 and 256 exports to w8a32.
Widescreen EdgePad: a light fine-tune
One checkpoint (SHA-256 2dc8fddcef9ef22f45b6a1573b9cf66de0f39b1b5172e4284dd693790bd6309e), trained from ZipDepth's Standard checkpoint in two short stages. Only the decoder was trained: 154,005 of 6.1 million parameters. The encoder is ZipDepth's.
- Stage 1 (5,000 steps, 170 s on an RTX 3090): shapes alternated between 384x384, 256x256, 224x224 and 512x288, so one set of weights serves every size.
- Data: 2,000 frames from 25 videos: 10 gameplay and YouTube, 8 film, live-action or CG, 7 animation. The frames came as 250 clips of 8 consecutive frames, with scene cuts marked so no temporal loss crosses one.
- Labels: the original square EdgePad-384's own output on every frame (an anchor, so the model doesn't drift from ZipDepth); Video Depth Anything Small (Apache-2.0) for frame-to-frame consistency on every clip; and Marigold V2 (Apache-2.0) on one hard keyframe per clip for spatial detail.
- A texture-leakage term at weight 0.5 discourages copying image texture into the depth.
- Temporal behaviour is learned only through the losses: the model has no memory, and each frame is still processed on its own.
- Stage 2 (1,000 steps at learning rate 5e-5, 44 s): decoder only again, for the exact 16:9 sizes.
The training frames aren't distributed.
Results, on held-out video against the untouched ZipDepth export at the same size (negative is better):
| Model | Compared with | Flicker on static scenes | Flow-warped error | Frame-to-frame change |
|---|---|---|---|---|
| 512x288 | stock 512x288 | -8.6% | -5.8% | -8.4% |
| 320x180 | stock 224x224 | -42.4% | -43.9% | -39.2% |
The cost is texture leakage: about 11% higher than the stock 224x224 at the 320x180 size. The 320x180 model is also visibly less detailed than the 384x384 one, as its size implies.
Exports: the TFLite models match their ONNX source with correlation 0.99999 (512x288 and 320x192), and LiteRT's GPU compatibility check reports no warnings. Meteor's ONNX files are the same weights with the 2x2 unpacking and clamp added to the graph and the weights embedded. The 672x384 file is an export of the same checkpoint at a larger size; it wasn't part of the held-out evaluation above.
Credits and licences
- ZipDepth: Fabio Tosi, MIT. These models are derivatives of its weights; its copyright notice is kept in
LICENSE. - Teacher models used only to make training labels, not included: Video Depth Anything Small (Apache-2.0) and Marigold V2 (Apache-2.0; built on Qwen-Image-Edit-2509, Apache-2.0).
- Fine-tuning, graph optimisation and exports: Nightfall's model research.