Machine learning supervised algorithms for gas hydrate identification and saturation estimation in marine reservoirs using well log data: A case study of NGHP-01-19B
Yi-fan Wu , Zheng Su , Takeshi Tsuji , Dai-dai Wu , Guang-rong Jin , Chao Yang , Chuang-ji Feng , Neng-you Wu
China Geology ›› 2026, Vol. 9 ›› Issue (3) : 519 -534.
Gas hydrates are increasingly recognized as a significant unconventional energy resource and a key factor in marine geohazards and the global carbon cycle. However, accurately identifying and quantifying hydrate-bearing formations remains challenging due to complex geophysical signatures and heterogeneous distribution. This study evaluates twelve supervised machine learning (ML) algorithms for two key tasks: Classification of hydrate-bearing layers and regression-based estimation of hydrate saturation, using well log and pore-water geochemical data from Site NGHP-01-19B. Two physically independent labeling frameworks are employed: One based on Archie’s law using resistivity (1350 samples, 29% hydrate-bearing), and another based on a three-phase velocity model (890 samples, 25% hydrate-bearing). A diverse set of models, including tree-based ensembles (Decision Tree, Random Forest, GBDT, XGBoost, LightGBM, CatBoost, Bagging, AdaBoost), kernel methods (SVM, SVR), instance-based learning (KNN), neural networks (MLP), and Gaussian Process models (GPR, GPC), are systematically compared using cross-validation and grid search. Ensemble methods consistently performed best in classification, with AdaBoost and GBDT, achieving test accuracies above 0.94 (Archie) and 0.98 (velocity-based). For regression, GPR delivered the most accurate hydrate saturation estimates (R2 > 0.99), while GBDT and Random Forest provided a strong balance of accuracy and computational efficiency. Notably, depth below seafloor (TDEP), though not a direct geophysical input, significantly enhanced model performance by acting as a proxy for stratigraphic and thermodynamic conditions. Group-based validation confirmed that random-sample splitting overestimates performance due to depth-wise autocorrelation, highlighting the importance of geologically informed model assessment. Overall, the consistent performance of ML models across both labeling schemes and input feature sets underscores their robustness and transferability, supporting their use as a reliable toolset for offshore gas hydrate reservoir characterization.
Gas hydrate / Machine learning algorithm / Classification / Regression / Well log data / Archie’s law / Marine geohazards / Global carbon cycle
| [1] |
|
| [2] |
|
| [3] |
|
| [4] |
|
| [5] |
|
| [6] |
|
| [7] |
|
| [8] |
|
| [9] |
|
| [10] |
|
| [11] |
|
| [12] |
|
| [13] |
|
| [14] |
|
| [15] |
|
| [16] |
|
| [17] |
|
| [18] |
|
| [19] |
|
| [20] |
|
| [21] |
|
| [22] |
|
| [23] |
|
| [24] |
|
| [25] |
|
| [26] |
|
| [27] |
|
| [28] |
|
| [29] |
|
| [30] |
|
| [31] |
|
| [32] |
|
| [33] |
|
| [34] |
|
| [35] |
|
| [36] |
|
| [37] |
|
| [38] |
|
| [39] |
|
| [40] |
|
| [41] |
|
| [42] |
|
| [43] |
|
| [44] |
|
| [45] |
|
| [46] |
|
| [47] |
|
| [48] |
|
| [49] |
|
| [50] |
|
| [51] |
|
| [52] |
|
| [53] |
|
| [54] |
|
| [55] |
|
| [56] |
|
| [57] |
|
| [58] |
|
| [59] |
|
| [60] |
|
| [61] |
|
| [62] |
|
| [63] |
|
| [64] |
|
| [65] |
|
| [66] |
|
| [67] |
|
| [68] |
|
| [69] |
|
| [70] |
|
| [71] |
|
| [72] |
|
| [73] |
|
/
| 〈 |
|
〉 |