在混合溫帶森林以多模態遙感觀測與機器學習估算地上生物量Aboveground biomass estimation using multimodal remote sensing observations and machine learning in mixed temperate forest
Lamahewage, S. H. G.; Witharana, C.; Riemann, R.; Fahey, R.; Worthley, T.|Scientific Reports 15: 31120|DOI: 10.1038/s41598-025-15585-6
狀態:AI_DRAFT_FROM_REVIEW|分級:A|閱讀深度:FULL_TEXT_CHECKED|Jacky 審核:False
森林數位孿生方法平台
aboveground biomassmultimodal remote sensingforest carbonsmall training datarandom forestLiDARSentinel-2NAIPFIAhyperparameter tuningvariable selectionUSAConnecticuttemperate forest
專討核心文獻定位
使用警示
本頁是文獻知識庫卡片,不等於可直接引用的最終查核稿。只有狀態升級為 CITABLE 後,才可直接進入論文引用候選。
為什麼納入這篇
This paper is a Ch6 core source because it is a concrete multimodal fusion case study that combines airborne LiDAR, Sentinel-2, and NAIP imagery with national forest inventory data through random forest, and it honestly reports the size of the fusion gain and the constraints of working with restricted, small training data.
結構式摘要|中英文對照
| 研究問題 | 在訓練資料有限且 FIA 樣區座標受政策限制的條件下,把光達與光學影像等多模態遙感資料和 FIA 樣區資料整合並對隨機森林做變數篩選與超參數調校,能否顯著提升大尺度森林地上生物量估算的準確度與穩健度。 Under limited training data and policy-restricted FIA plot coordinates, can integrating multimodal remote sensing such as LiDAR and optical imagery with FIA plot data and applying variable selection and hyperparameter tuning to random forest significantly improve the accuracy and robustness of large-area forest aboveground biomass estimation. |
|---|---|
| 資料來源 | 研究區為美國康乃狄克州,以橡樹與山核桃為主的溫帶森林。使用 2016 年州級機載光達(點密度 2 點每平方公尺)、2016 年 NAIP 航照(0.6 公尺)、Sentinel-2 影像、土壤與森林覆蓋圖,搭配 2016 年 42 個 FIA 真實樣區共 168 個子樣區,剔除離群後留 142 個子樣區,並萃取 67 個解釋變數,生物量響應值取自 FIA 以成分比例法計算的子樣區地上生物量。 The study area is the state of Connecticut, USA, an oak- and hickory-dominated temperate forest. It uses 2016 statewide airborne LiDAR (2 points per square meter), 2016 NAIP imagery (0.6 m), Sentinel-2 imagery, and soil and forest-cover maps, together with 42 true FIA plots and 168 subplots from 2016 reduced to 142 subplots after outlier removal, with 67 explanatory variables and an AGB response derived from FIA component-ratio-method estimates. |
| 方法 | 把 67 個變數依來源分為三組,第一組全部 67 個變數(光達加影像)、第二組僅光達、第三組僅影像。各組以 IncNodePurity 與 5 摺交叉驗證 R2 做前向變數篩選,再經預設、ntree 調整與格點搜尋三種設定的超參數調校,共建立九個隨機森林模型,並以 80 比 20 切分、保留 20% 做測試,最後與三個既有生物量地圖產品比較。 The 67 variables are split by source into Group-1 all 67 variables (LiDAR plus image), Group-2 LiDAR only, and Group-3 image only. Each group undergoes forward variable selection using IncNodePurity and 5-fold cross-validated R2, then hyperparameter tuning across default, ntree-adjusted, and grid-search settings, yielding nine random forest models with an 80:20 train-test split and a held-out 20% test set, finally compared against three existing biomass map products. |
| 主要結果 | 最佳模型為第一組格點調校的 All_RF03,採用前 28 個解釋變數,測試 RMSE 27.19 Mg/ha、R2 0.41、可解釋變異 40.54%;僅光達的 LiDAR_RF03 為 RMSE 27.88 Mg/ha、R2 0.23,僅影像的 Img_RF03 為 RMSE 26.50 Mg/ha、R2 0.31。結論段指出加入影像資料相較單用光達使 R2 提升 34.78%;第一組最重要的 28 個變數中有 68% 來自光達高度相關變數。在與既有地圖產品比較中,本研究模型 RMSE 28.33 Mg/ha 明顯優於 90 公尺、250 公尺與國家尺度地圖(RMSE 分別為 145.35、82.47、81.84 Mg/ha)。 The best model is the Group-1 grid-tuned All_RF03 using the top 28 variables, with test RMSE 27.19 Mg/ha, R2 0.41, and 40.54% variance explained; LiDAR-only LiDAR_RF03 is RMSE 27.88 Mg/ha and R2 0.23, image-only Img_RF03 is RMSE 26.50 Mg/ha and R2 0.31. The conclusion states that adding image data improved R2 by 34.78% relative to LiDAR alone; 68% of the top 28 variables in Group-1 are LiDAR height-related. In the map-product comparison the study model RMSE 28.33 Mg/ha clearly outperformed the 90 m, 250 m, and national-scale maps (RMSE 145.35, 82.47, 81.84 Mg/ha respectively). |
| 限制 | 資料量小,僅 142 個子樣區,且因 FIA 不同意政策只能在離站電腦使用單年 2016 資料,未加入生長增量;FIA 真實座標仍有約 8 公尺定位誤差,為遙感變數帶入未知雜訊;模型對 100 Mg/ha 以上的高生物量區有低估傾向;森林型別變數重要性偏低,無法分層分析不同森林型別的表現;整體 R2 0.41 屬於中等,作者自承主因是訓練資料受限。 The dataset is small at only 142 subplots, and because of FIA non-consent policy the analysis used single-year 2016 data on off-site computers without growth increments; true FIA coordinates still carry about 8 m geolocation error that adds unknown noise to RS variables; the model tends to underestimate high-biomass areas above 100 Mg/ha; forest-type variable importance was low, preventing stratified analysis by forest type; the overall R2 of 0.41 is moderate, which the authors attribute mainly to restricted training data. |
Key Findings
| 發現 | 證據 | 確定性 |
|---|---|---|
| Combining LiDAR with optical image metrics through random forest improves AGB prediction accuracy over LiDAR alone, but the gain is moderate rather than dramatic. | Conclusion states combining RS image data with LiDAR metrics improved R2 by 34.78% compared with LiDAR alone; the optimal All_RF03 reached RMSE 27.19 Mg/ha and R2 0.41 versus LiDAR-only R2 0.23. | checked_against_original_txt |
| LiDAR height-related variables dominate the most important predictors in the combined model. | Results and discussion report that 68% of the top 28 selected variables in Group-1 were LiDAR-derived, with the 95th height percentile the single most important variable. | checked_against_original_txt |
| Careful hyperparameter tuning and variable selection can rescue useful AGB models even from small, policy-constrained training datasets. | Grid tuning raised Group-1 R2 from 0.19 to about 0.40-0.41 with a 23.71% RMSE reduction; the discussion frames this as the value of tuning under N=142. | checked_against_original_txt |
Key Figures and Tables
公開網站原則:未確認授權前,不直接複製原文圖表;優先使用自製圖表導讀或重繪圖。
| 項目 | 內容 | 關鍵數字 | Jacky 判讀 | 重用策略 |
|---|---|---|---|---|
| Table 1, Table 2, Table 3, Fig. 4 | Table 1 lists the hyperparameter combinations for the nine models across three groups; Table 2 gives the test evaluation metrics for all nine RF models; Table 3 compares the optimal RF model against three existing AGB map products; Fig. 4 shows the mean 5-fold CV R2 as variables are added by IncNodePurity rank for each group. | Best All_RF03: RMSE 27.19 Mg/ha, R2 0.41, 40.54% variance explained, 28 variables. LiDAR_RF03 R2 0.23, Img_RF03 R2 0.31. Map comparison RMSE: RF 28.33 vs Map_1 145.35, Map_2 82.47, Map_3 81.84 Mg/ha. 67 candidate variables; 142 subplots; nine RF models. | Use Table 2 and Table 3 to argue that a tuned multimodal RF can beat coarse-resolution map products at plot scale, then position this as a middle-layer AGB engine that could feed a Taiwan forest digital twin. | License is CC BY-NC-ND, so do not adapt the original figures; redraw a self-made comparison chart and cite the source. |
Extracted Evidence Table
| 可支撐主張 | 指標或結果 | 原文位置 | 可引用 | 備註 |
|---|---|---|---|---|
| A tuned multimodal random forest model outperforms coarse-resolution AGB map products at FIA subplot scale. | Optimal RF achieved RMSE 28.33 Mg/ha and R2 0.41, versus Map_1 (90 m) RMSE 145.35 R2 0.01, Map_2 (250 m) RMSE 82.47 R2 0.20, Map_3 (national) RMSE 81.84 R2 0.10. | Comparison with available map products section and Table 3, p.9, p.12. | True | Report alongside the resolution mismatch caveat the authors raise, since coarse maps are not designed for plot-level comparison. |
| Adding optical image metrics to LiDAR improves AGB model fit by a moderate margin. | Conclusion reports R2 improved by 34.78% when combining RS image data with LiDAR metrics versus LiDAR alone; discussion states the combination increased overall R2-based performance by about 24%. | Discussion p.10 and Conclusion p.13. | True | The paper gives two framings of the gain (24% in discussion, 34.78% in conclusion); cite the conclusion figure but note the discussion phrasing to stay precise. |
| LiDAR height metrics are the dominant predictors of temperate-forest AGB in this multimodal model. | 68% of the top 28 selected variables were LiDAR-derived; the 95th height percentile was the most important variable in both the combined and LiDAR-only groups. | Variable importance section, p.7, and Discussion p.10-11. | True | Use to support height as the primary structural driver, with image metrics as complementary heterogeneity information. |
Critical Appraisal
Strengths
- Concrete, transparent multimodal fusion case study integrating LiDAR, Sentinel-2, NAIP, and FIA inventory data.
- Honest reporting of a moderate R2 with explicit discussion of small-sample and data-policy constraints.
- Systematic three-group variable design plus tuning that isolates the contribution of each data source.
Weaknesses
- Small training set (N=142) and single-year 2016 data limit generalizability.
- Underestimates high-biomass plots above 100 Mg/ha, so it is weak exactly where carbon stocks are largest.
- Strongly tied to US FIA infrastructure and Connecticut conditions, so transfer needs local recalibration.
| Validation quality | moderate; held-out 20% testing and map-product comparison are sound but accuracy is constrained by data size and FIA geolocation error |
|---|---|
| Transferability to Taiwan | medium; the multimodal RF workflow transfers well as a method, but the FIA data pipeline and oak-hickory temperate setting differ from Taiwan and would need local inventory and allometric calibration |
| Risk of overclaiming | Do not present R2 0.41 as high accuracy or claim multimodal fusion guarantees large gains; the paper itself frames the result as moderate and data-limited. |
與 Jacky 博論 / Review 的用途
| 博士論文 | Supports the dissertation's middle-layer argument that multimodal remote sensing plus machine learning can produce fine-resolution AGB surfaces, while honestly bounding how much fusion and tuning help under limited ground data. |
|---|---|
| TJFS Review | Gives the TJFS review a recent, citable temperate-forest example of LiDAR-plus-optical fusion with quantified accuracy and a candid limitations section, useful for the multi-source fusion chapter. |
| 可引用句候選 | 2025 年,Lamahewage 等人發表的文獻中指出,在康乃狄克州溫帶森林結合光達與光學影像並對隨機森林做超參數調校後,最佳模型測試 RMSE 為 27.19 Mg/ha、R2 為 0.41,且加入影像資料相較單用光達使 R2 提升 34.78%。 |
| 不可用來主張 | Do not cite this paper as evidence of high-accuracy or operational large-scale AGB mapping, nor as proof that multimodal fusion always yields large gains. |
授權與圖表重用
| Article license | CC BY-NC-ND 4.0 |
|---|---|
| Figure reuse policy | DO_NOT_ADAPT_OR_REDISTRIBUTE_ADAPTED_FIGURES_NONCOMMERCIAL_ND_REDRAW_INSTEAD |
| Notes | Open Access, The Author(s) 2025, Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License (http://creativecommons.org/licenses/by-nc-nd/4.0/). The ND clause means adapted versions of the original figures may not be shared; redraw self-made diagrams and only reuse unmodified figures non-commercially with attribution. |
待查核清單
- Confirm whether a Chinese translation_md, txt, and seminar導讀 pptx exist for this paper; currently marked 待查.
- Reconcile the two reported fusion-gain figures (24% in discussion vs 34.78% in conclusion) before quoting a single number.
- Because the licence is CC BY-NC-ND, prepare a redrawn comparison figure rather than adapting Table 3 or Fig. 4 directly.