Depth Anything Rethought for Tiny Models

DepthART

ACMMM 2026

Scaling Foundation Monocular Depth to Tiny Models

New · September 2026

DepthART now supports MobileNetV4 backbones for earlier-generation, resource-constrained devices. MobileNetV4-M-slim-SPF pairs a slimmer encoder with Single-Path Pyramid Fusion for portable BPU/NPU deployment.

6.0M / 8.0MRelative-S / Metric-S parameters
0.92 msDepthART-S 224 · A6000 TensorRT FP16
1088.9DepthART-S 224 · A6000 FPS · TensorRT FP16
247.8 FPSOrin NX · TensorRT FP16
0.971 δ1DepthART-L · NYUD v2 zero-shot
Videos from ADVIO datasets, captured on iPhone

Relative Depth in the Wild

DepthART processes handheld portrait video while preserving fine boundaries and accurate depth.

RGB
DepthART-L · Relative
Algorithm Overview

DepthART

Recent geometric foundation models have advanced monocular depth estimation, yet their benefits remain limited for tiny models. We present DepthART, a compact model designed for robust on-device depth estimation across diverse scenes. To address dataset-specific overfitting and unstable metric adaptation under camera shifts, DepthART combines bias-resistant data sampling with camera-conditioned fine-tuning that preserves the distilled encoder while adapting metric scale using camera intrinsics. These designs improve both cross-dataset generalization and metric depth prediction in capacity-constrained models.

1
Bias-resistant distillationRebalance a 44M multi-source corpus, then distill DepthAnything v2-L.
2
Camera-conditioned fine-tuningFreeze the trunk and adapt scale using camera prompts and a multi-query head.
Multi-platform Deployment

Accuracy vs Speed

Platform
Power mode
Compute precision
Depth task
Dataset
RTX A6000PyTorch FP32Relative · NYUD v2
Loading benchmark data…
0.9000.9250.9500.9751.000Model latency (ms, log scale) →NYUD v2 δ1 ↑
Additional comparison familiesAwaiting aligned A6000 FP32 model-only latency
DepthART encoderTinyViM-STinyViM-BTinyViM-LMobileNetV4-SMobileNetV4-MMobileNetV4-M-slim-SPF
Move across the plot to magnify nearby latency values.Comparison methods use affine-invariant δ1.
Robot deployment

Metric Point Cloud

DepthART-Metric-LFine-tuned on NYUD v2 only
Cafe · Scene 1Metric 3D reconstruction
Cite this work

DepthART

Accepted to ACM Multimedia 2026. The final citation and public model links will be added with the camera-ready release.

@inproceedings{depthart2026,
  title = {DepthART: Scaling Foundation Monocular Depth to Tiny Models},
  author = {Feng Xue and Wu Chen and Mingshuai Zhao and Guofeng Zhong and Anlong Ming and Haozhe Wang and Dianqiao Lei and Zhaowen Lin and Haiyang Zhang and Nicu Sebe},
  booktitle = {ACM Multimedia},
  year = {2026}
}