森林地上生物量估算中的機器學習與多源遙感之回顧Machine Learning and Multi-source Remote Sensing in Forest Aboveground Biomass Estimation: A Review
Nguyen A, Saha S|arXiv preprint (NeurIPS 2025 Workshop, Tackling Climate Change with Machine Learning) arXiv:2411.17624|DOI: 10.48550/arXiv.2411.17624
狀態:AI_DRAFT_FROM_REVIEW|分級:A|閱讀深度:FULL_TEXT_CHECKED|Jacky 審核:False
森林數位孿生方法平台
AGBmachine learningmulti-source remote sensingsystematic reviewforest carbonrandom forestXGBoostfeature selectionSentinel-1 Sentinel-2 LiDAR fusionsystematic literature reviewglobal
專討核心文獻定位
使用警示
本頁是文獻知識庫卡片,不等於可直接引用的最終查核稿。只有狀態升級為 CITABLE 後,才可直接進入論文引用候選。
為什麼納入這篇
This paper is a recent systematic review that quantitatively summarizes which machine-learning methods and which multi-source remote-sensing combinations are most used and most effective for forest AGB estimation, providing a strong synthesis anchor for the Ch5 machine-learning narrative.
結構式摘要|中英文對照
| 研究問題 | 在結合機器學習與多源遙感估算森林地上生物量時,哪些機器學習方法、哪些遙感資料來源與組合最常被使用且表現最好,並且如何把這些發現用容易理解的方式傳達給非專業讀者。 When combining machine learning and multi-source remote sensing for forest aboveground biomass estimation, which ML methods and which RS data sources and combinations are most frequently used and most effective, and how can these findings be communicated accessibly to non-expert readers. |
|---|---|
| 資料來源 | 從超過 80 篇相關研究中,依五項嚴格納入條件系統性篩選出 25 篇文獻進行量化分析。納入條件包含全文可取得、使用機器學習、使用多種遙感資料來源、研究目標為估算森林碳量(AGB、BGB 或土壤碳)、發表於近十年(2014 至 2024 年)。對每篇研究蒐集遙感來源、機器學習方法、最佳方法、限制與未來方向、最終任務、研究地點、森林類型與研究尺度等資訊。 From over 80 related studies, 25 papers were systematically selected for quantitative analysis under five strict inclusion criteria: full paper accessible, used ML, used multiple RS sources, end goal was estimating forest carbon (AGB, BGB, or soil carbon), and published in the recent ten years (2014-2024). For each study the authors collected RS sources, ML methods, best-performing method, limitations and future steps, ultimate task, study location, forest type, and study scale. |
| 方法 | 以系統性文獻回顧方法,先用設定的搜尋詞透過 Google Scholar 內部 API 檢索建庫,依五項納入條件取前 25 篇。對 25 篇逐篇編碼遙感來源、機器學習方法、是否做多模型比較與最佳模型、限制、最終任務、地點、森林類型與尺度,並用 Pandas、Matplotlib、Seaborn、NumPy 做資料整理與視覺化,產出頻率圖、二元共現熱圖與彙整表,另建一個可篩選的互動資料庫。 Using a systematic literature review approach, the authors retrieved and built a database with defined search terms via a Google Scholar internal API, then took the first 25 papers meeting five inclusion criteria. They coded each of the 25 papers for RS sources, ML methods, multi-model comparison and best model, limitations, ultimate task, location, forest type, and scale, and used Pandas, Matplotlib, Seaborn, and NumPy to organize and visualize the data, producing frequency charts, a binary co-occurrence heatmap, and a summary table, plus an interactive filterable database. |
| 主要結果 | 隨機森林出現於約 88% 的研究,為最常被使用的方法;25 篇中有 11 篇做了多模型比較,RF 雖出現於所有比較研究卻僅在 4 篇被認定最佳,XGBoost 雖只用於 4 篇卻在其中 3 篇被認定最佳。最常用的遙感來源依序為 Sentinel-1、Sentinel-2、ALOS-PALSAR、Landsat、MODIS、GLAS/ICESat LiDAR 與 GEDI LiDAR;Sentinel-1、Sentinel-2 與太空載 LiDAR 最常一起使用。特徵選擇被多篇研究認定為影響模型表現的關鍵因素,物候特徵與來自 DEM 或 SRTM 的三維結構資訊在缺乏 LiDAR 時尤其重要。地理上中國約占所有研究地點的一半。 Random Forest appeared in about 88% of studies, the most frequently used method; 11 of 25 papers compared multiple models, and although RF appeared in all comparison studies it was found best in only 4, whereas XGBoost was used in only 4 studies but found best in 3 of them. The most-used RS sources, in order, were Sentinel-1, Sentinel-2, ALOS-PALSAR, Landsat, MODIS, GLAS/ICESat LiDAR, and GEDI LiDAR; Sentinel-1, Sentinel-2, and spaceborne LiDAR were most often used together. Feature selection was identified by many studies as a key factor in model performance, and phenological features and 3D structural information from DEM or SRTM were especially important when LiDAR was unavailable. Geographically, China accounted for about half of the study locations. |
| 限制 | 作者坦言為大學生且幾乎無期刊訂閱,第三項以全文可取得為納入條件,使樣本偏向開放取用文獻,可能有取樣偏差;僅取搜尋結果前 25 篇且樣本量小,難以做嚴謹統計推論;對最佳方法的判定依賴各研究摘要或結論中是否出現最佳或最高等關鍵詞,標準不一。森林類型分類為作者主觀整理且非互斥,地理上高度集中於中國,外推性受限。 The authors acknowledge they are college students with almost no journal subscriptions, so the inclusion criterion of full-text accessibility biases the sample toward open-access papers, introducing possible selection bias. They only took the first 25 search results and the small sample size limits rigorous statistical inference. Best-method determination relied on whether keywords like best or highest appeared in each study's abstract or conclusion, an inconsistent standard. Forest-type categories were subjective and non-exclusive, and the strong geographic concentration in China limits generalizability. |
Key Findings
| 發現 | 證據 | 確定性 |
|---|---|---|
| Random Forest was the most frequently used ML method, appearing in about 88% of the reviewed studies. | Original Abstract and Results: Random Forest was used in around 88% of the studies as the model for AGB estimation and sometimes for intermediate tasks (p.1 Abstract, p.3 Results). | checked_against_original_txt |
| XGBoost showed superior performance more often than RF when methods were compared head-to-head. | Original Results: of 11 studies comparing multiple ML methods, RF was found best in only 4, while XGB was used in only 4 studies but found best in 3; Abstract states XGB showed superior performance in 75% of studies in which it was compared (p.1, p.3). | checked_against_original_txt |
| Sentinel-1 was the most utilized RS source, and multi-sensor combinations of optical, radar, and spaceborne LiDAR proved especially effective. | Original Abstract and Results: Sentinel-1 emerged as the most utilized source, followed by Sentinel-2, ALOS-PALSAR, Landsat, MODIS; Sentinel-1, Sentinel-2 and spaceborne LiDAR were most often used together (p.1, p.3-4). | checked_against_original_txt |
| Feature selection, phenological variables, and 3D structural data (DEM/SRTM) were critical accuracy drivers, especially when LiDAR was unavailable. | Original Results & Discussion: feature selection was a critical factor; phenological variables raised R-squared and enabled time-consistent AGB models; DEM/SRTM was often the most important predictor without LiDAR (p.3-4). | checked_against_original_txt |
Key Figures and Tables
公開網站原則:未確認授權前,不直接複製原文圖表;優先使用自製圖表導讀或重繪圖。
| 項目 | 內容 | 關鍵數字 | Jacky 判讀 | 重用策略 |
|---|---|---|---|---|
| Table 1 | Compares passive optical, active optical (LiDAR), and radar (SAR) sensors by capability and limitation, listing Sentinel-2, GLAS/ICESat, Sentinel-1, MODIS, ALOS-PALSAR, Landsat. | Three sensor modality categories; passive optical (Sentinel-2, MODIS, Landsat) widely and freely available but cannot work at night or penetrate canopy; LiDAR measures 3D structure but is expensive; radar (Sentinel-1, ALOS-PALSAR) penetrates clouds but has saturation issues. | 可作為多源遙感互補性的入門對照表,說明為何光學加雷達加 LiDAR 的組合最有效。 | Redraw a simplified sensor-comparison table after confirming the arXiv license; do not reproduce the original table publicly until license checked. |
| Figure 2 | Heatmap showing which pairs of remote-sensing sources were most often used together across the reviewed studies. | Sentinel-1, Sentinel-2, and spaceborne LiDAR (GEDI) most often used together; without LiDAR, common pairs were Landsat+Sentinel-1, MODIS+Sentinel-2, MODIS+PALSAR. | 顯示主流多源融合的實際搭配模式,可作為討論感測器組合選擇的依據。 | Self-draw a schematic co-occurrence summary; do not reproduce the original heatmap publicly until license checked. |
| Table 2 | Per-study listing of remote-sensing data sources and ML methods for the 25 reviewed studies. | 25 studies coded; data-source combinations and ML method sets (RF, XGB, SVM, CNN, LSTM, kNN, CatBoost, LGBM, GBM and others) listed per study. | 可作為快速查找各研究所用感測器與模型搭配的索引表。 | Extract only the structured data needed for synthesis; do not reproduce the full original table publicly until license checked. |
Extracted Evidence Table
| 可支撐主張 | 指標或結果 | 原文位置 | 可引用 | 備註 |
|---|---|---|---|---|
| Random Forest is the dominant ML method but is not always the best performer when directly compared. | RF used in ~88% of studies; in 11 head-to-head comparisons RF best in only 4, while XGB best in 3 of the 4 studies it appeared in (XGB superior in 75% of comparisons it was in). | Abstract (p.1); Results & Discussions (p.3). | True | Frequencies are over a small systematic sample of 25 papers; cite as a review synthesis, not population statistics. |
| Multi-sensor fusion of optical, radar, and spaceborne LiDAR is the most effective RS strategy, with Sentinel-1 most utilized. | Sentinel-1 most-used source; Sentinel-1+Sentinel-2+spaceborne LiDAR most common combination; LiDAR recommended to overcome optical/radar saturation. | Abstract (p.1); Results & Discussions (p.3-4); Conclusion (p.4). | True | Confirmed against original abstract, results, and conclusion sections. |
| Feature selection, phenology, and 3D structure data materially improve AGB model accuracy. | Lasso reduced 30+ parameters to seven; phenological variables raised R-squared and enabled time-consistent models; DEM/SRTM often the most important predictor without LiDAR. | Results & Discussions (p.3-4). | True | Specific per-study examples (Li et al., Huang et al.) drawn from the review's own citations; attribute to the review's synthesis. |
Critical Appraisal
Strengths
- 明確的系統性回顧方法與五項納入條件,並把機器學習方法與多源遙感組合的使用頻率與表現量化呈現。
- 同時整理感測器互補性、特徵選擇、物候與三維結構資料等實務洞見,並提供可篩選的互動資料庫。
- 結論給出明確且可操作的建議,例如至少結合被動光學與雷達、缺 LiDAR 時納入 DEM、不同階段用不同模型。
Weaknesses
- 樣本僅 25 篇且以全文可取得為條件,偏向開放取用文獻,存在取樣偏差。
- 最佳方法判定僅依摘要或結論中的關鍵詞,標準主觀;地理上高度集中於中國,外推性受限。
- 為工作坊短篇 preprint 而非完整期刊論文,統計推論深度有限,森林類型分類為作者主觀整理。
| Validation quality | review synthesis over 25 systematically selected papers; no independent empirical validation, frequencies are descriptive over a small sample |
|---|---|
| Transferability to Taiwan | moderate to high as a method-selection reference; the multi-sensor fusion and feature-selection recommendations transfer well, though the China-heavy sample and arid/temperate species differ from Taiwan subtropical forests |
| Risk of overclaiming | Do not present the 88% RF or 75% XGB figures as population-level statistics; they are frequencies over a small 25-paper sample with open-access selection bias. |
與 Jacky 博論 / Review 的用途
| 博士論文 | 支撐博論在機器學習估算 AGB 的方法選擇,提供 RF 作為基線、XGBoost 與 CNN 作為進階選項,以及多源遙感融合與特徵選擇的綜整依據。 |
|---|---|
| TJFS Review | 在 TJFS review 的 Ch5 機器學習章節可作為近期系統性回顧的綜整錨點,量化呈現主流模型與感測器組合的使用頻率與表現。 |
| 可引用句候選 | 2024 年,Nguyen 等人發表的文獻中指出,系統性回顧 25 篇結合機器學習與多源遙感的森林生物量研究後,隨機森林雖最常被使用,但極限梯度提升在多模型比較中往往表現最佳,且 Sentinel-1 與光學及太空載 LiDAR 的多感測器融合最為有效。 |
| 不可用來主張 | 不要把 88% 或 75% 等頻率當成母體統計值,也不要以本篇單獨宣稱某模型在所有森林環境中皆最佳,因其樣本僅 25 篇且偏向開放取用文獻。 |
授權與圖表重用
| Article license | UNKNOWN |
|---|---|
| Figure reuse policy | DO_NOT_REUSE_ORIGINAL_FIGURES_PUBLICLY_UNTIL_LICENSE_CHECKED |
| Notes | arXiv preprint(arXiv:2411.17624v2,NeurIPS 2025 Workshop)。原文 PDF 文字未見明確 CC 授權聲明,預設為 arXiv 非專屬授權,圖件改繪或重用前須回原文授權頁確認,待查。 |
待查核清單
- 回 arXiv 原文授權頁確認 article_license(CC-BY 或 arXiv 非專屬),再決定圖表呈現方式與是否可填 full_translation。
- 如需引用各模型使用頻率與最佳模型次數,回 Figure 1、Figure 3、Figure 4 與 Table 2 核對精確數字。
- 注意此篇 arXiv:2411.17624 為 NeurIPS 2025 Workshop 短篇 preprint,與跨產業 Nguyen 2025 Sensors 製造數位孿生為不同篇,引用時勿混淆。