HU Qiao, JIANG Weiguo, ZHAO Zongying, KANG Yu, LI Jiating, ZHANG Jingjing, YU Jiahua, LI Dongxue. Transferring Prior Visual Knowledge in Large-scale Convolutional Neural Networks Facilitates Interpretable and Multi-scale Wetland Mapping. Chinese Geographical Science. DOI: 10.1007/s11769-026-1699-2
Citation: HU Qiao, JIANG Weiguo, ZHAO Zongying, KANG Yu, LI Jiating, ZHANG Jingjing, YU Jiahua, LI Dongxue. Transferring Prior Visual Knowledge in Large-scale Convolutional Neural Networks Facilitates Interpretable and Multi-scale Wetland Mapping. Chinese Geographical Science. DOI: 10.1007/s11769-026-1699-2

Transferring Prior Visual Knowledge in Large-scale Convolutional Neural Networks Facilitates Interpretable and Multi-scale Wetland Mapping

  • Large-scale Convolutional Neural Networks (CNNs) contain rich visual semantics that are generalizable across diverse visual objects but are challenging to adapt to downstream natural-system mapping due to their time-consuming and opaque inference processes. This study presents KB-Net (Knowledge Base Net), which leverages large-scale CNNs as prior visual knowledge bases to enable cost-effective and interpretable land-cover mapping. The algorithm employs an expert system to heuristically retrieve and integrate optimal visual semantics from multiple spatial scales from CNN-based knowledge bases, facilitating multi-scale land cover mappings. Furthermore, the expert system mimics human-like reasoning by incorporating physically explicit distance metrics to produce visually interpretable inferences. KB-Net was evaluated on multi-temporal wetland datasets under varying data forms, including unmanned aerial vehicle (UAV) and satellite, visual semantics, and scale configurations. UAV-based experiments demonstrated that KB-Net, using visual semantics derived from pre-trained large-scale CNNs, achieved high segmentation accuracies (0.720–0.820), comparable to locally trained CNNs (0.760–0.850), and provided a 2%–7% improvement in Intersection over Union (IoU) over raw image features. Satellite-based experiments further showed that KB-Net embeddings yielded higher overall accuracies (0.740–0.796) than direct satellite features (0.731–0.751), corresponding to an 11.7% IoU improvement. Comparisons with the state-of-the-art Segment Anything Model (SAM) indicated that KB-Net (0.440 IoU) outperformed SAM in both zero-shot (0.386 IoU) and few-shot modes (0.266 IoU), representing a 5.4%–17.4% IoU improvement. Because the prior visual knowledge used for visual semantics is readily deployable, KB-Net enables cost-effective, multi-scale land-cover mapping within minutes without extensive local labeling, training, or tuning. This study demonstrates a practical pathway to interpret and exploit prior visual knowledge in large-scale CNNs for downstream land-cover mapping, particularly in dynamic environmental contexts.
  • loading

Catalog

    /

    DownLoad:  Full-Size Img  PowerPoint
    Return
    Return