1
Bias-resistant distillationRebalance a 44M multi-source corpus, then distill DepthAnything v2-L.
Depth Anything Rethought for Tiny Models
Scaling Foundation Monocular Depth to Tiny Models
DepthART processes handheld portrait video while preserving fine boundaries and accurate depth.
Recent geometric foundation models have advanced monocular depth estimation, yet their benefits remain limited for tiny models. We present DepthART, a compact model designed for robust on-device depth estimation across diverse scenes. To address dataset-specific overfitting and unstable metric adaptation under camera shifts, DepthART combines bias-resistant data sampling with camera-conditioned fine-tuning that preserves the distilled encoder while adapting metric scale using camera intrinsics. These designs improve both cross-dataset generalization and metric depth prediction in capacity-constrained models.



MiDaS v3.1 · LeViTDepthART-LAccepted to ACM Multimedia 2026. The final citation and public model links will be added with the camera-ready release.
@inproceedings{depthart2026,
title = {DepthART: Scaling Foundation Monocular Depth to Tiny Models},
author = {Feng Xue and Wu Chen and Mingshuai Zhao and Guofeng Zhong and Anlong Ming and Haozhe Wang and Dianqiao Lei and Zhaowen Lin and Haiyang Zhang and Nicu Sebe},
booktitle = {ACM Multimedia},
year = {2026}
}