- Research Article
- 10.1109/jsen.2025.3611842
Self-Supervised Monocular Depth Estimation Leveraging Heterogeneous Network Dynamic Distillation and Cross Channel-Scale Attention
- Jan 01, 2026
- IEEE Sensors Journal
- Jiaojiao Fang
Self-supervised monocular depth estimation offers a promising alternative to costly manual annotations, yet it faces a critical challenge when test-time image resolution differs from training data, often resulting in depth maps with inconsistent structures and loss of fine details. Existing methods often overlook the importance of selecting optimally-performing teacher models for resolution adaptation, thereby limiting effective mutual learning and ignoring architectural compatibility across varying input resolutions. We propose an adaptive multi-resolution framework that dynamically selects the most suitable teacher model from heterogeneous depth estimation networks to enhance the robustness and stability of multi-space distillation under the guidance of scene structure and cross-resolution consistency. The proposed framework specifically incorporates three key components: firstly, heterogeneous network architectures utilizing sub-pixel convolution and cross-channel-scale attention mechanisms to reduce training instability caused by output noise across resolutions; secondly, a dynamic multi-space cross-resolution distillation strategy that enables adaptive mutual learning through teacher-student role-switching based on temporal prediction consistency; and thirdly, an additional structural regularization term within the mutual learning framework to enhance structure-aware detail representation in both low- and high-resolution pathways. Extensive evaluation on KITTI and Make3D benchmarks demonstrates that the proposed method not only achieves competitive performance compared to state-of-the-art self-supervised monocular depth estimation approaches, but also improves cross-resolution generalization by 1.2% in squared relative error on the KITTI dataset.
Read more