



































格式


    

 Academic Journal of Science, Engineering and Technology 

Vol.7, Issue 1; January - Febuary 2022; 

1252 Columbia Rd NW, Washington DC, United States 

https://topjournals.org/index.php/AJSET/index; mail: topacademicjournals@gmail.com 

  

 

 

1 | A c a d e m i c  J o u r n a l  o f  S c i e n c e ,  E n g i n e e r i n g  a n d  T e c h n o l o g y  

|  https://topjournals.org/index.php/AJSET 

A COMPREHENSIVE APPROACH TO RAILWAY LANDSLIDE RISK ASSESSMENT 

USING HIERARCHICAL GREY RELATION ANALYSIS 

 
1He, Xiao-Rong and 2Ji, Yu-Tong  
1Baoji Electric Service Section of Xi'an Railway Bureau, Baoji, 721000, China 
2Baojidong Railway Station of Xi'an Railway Bureau, Baoji, 721000, China 

 

Abstract: The West Sichuan Railway in China traverses challenging terrain characterized by steep slopes and a 

high seismic activity, leading to the frequent occurrence of landslides and other geological disasters during both 

construction and operation. Landslides pose a substantial threat to passenger safety and railway infrastructure. In 

the Chengdu Baiyu section alone, 126 landslides have been recorded, emphasizing the urgency of landslide risk 

assessment and prevention [1]. Landslide risk refers to the likelihood of slope failures evolving into various forms 

of disasters under the influence of multiple factors [3]. This study focuses on enhancing landslide risk assessment 

by optimizing the selection of evaluation factors. 

Previous research has explored various methods for factor selection and evaluation models in landslide risk 

assessment. One approach employed rough set analysis, correlation analysis, and principal component analysis 

to identify landslide evaluation factors, subsequently utilizing support vector machine models for landslide 

susceptibility evaluation [4]. This method demonstrated improvements in evaluation accuracy by reducing and 

refining evaluation factors. Another study applied the Apriori algorithm for correlation analysis of landslide 

evaluation factors and employed the Random Forest model for landslide susceptibility evaluation, resulting in 

more accurate evaluation results closely aligned with actual landslide distribution [5]. Additionally, genetic 

algorithms and rough set analysis were employed to screen landslide evaluation factors, coupled with BP neural 

networks for landslide susceptibility evaluation, revealing higher evaluation model accuracy with reduced factors 

[6]. Models like BP neural networks, support vector machines, Logistic regression, and Random Forest have 

consistently shown strong performance in landslide susceptibility assessment [7][8][9][10]. In a comparative 

study, the Random Forest model outperformed Logistic regression, Multilayer perceptron, and gradient 

enhancement tree models in landslide risk assessment [11]. 

However, a gap exists in the selection of evaluation factors as some studies have neglected to consider the 

contribution of these factors to landslide events, potentially compromising calculation efficiency and evaluation 

accuracy. Furthermore, the conventional data analysis method for factor screening often involves calculating the 

contribution rate of factors, which may introduce errors and outliers into the evaluation factor data extracted from 

landslide events. This could lead to the inadvertent elimination of important factors during data calculation and 

analysis, ultimately affecting the accuracy of evaluation results. Therefore, this study seeks to address these 

limitations by proposing a novel approach to evaluate landslide risk that incorporates a comprehensive 

consideration of evaluation factors and employs advanced data analysis techniques. 

Keywords: Landslide risk assessment, Evaluation factors, Factor selection, Data analysis, Random Forest model 

mailto:topacademicjournals@gmail.com


    

 Academic Journal of Science, Engineering and Technology 

Vol.7, Issue 1; January - Febuary 2022; 

1252 Columbia Rd NW, Washington DC, United States 

https://topjournals.org/index.php/AJSET/index; mail: topacademicjournals@gmail.com 

  

 

 

2 | A c a d e m i c  J o u r n a l  o f  S c i e n c e ,  E n g i n e e r i n g  a n d  T e c h n o l o g y  

|  https://topjournals.org/index.php/AJSET 

1. Introduction  

The terrain along the West Sichuan Railway in China is steep, with frequent occurrence of moderate to strong 

earthquakes. The geological environment along the railway is complex, and geological disasters such as 

landslides, collapses, and mudslides may occur during construction and operation [1]. Landslides pose a serious 

threat to the lives and property of passengers and railway lines. According to preliminary statistics, there are 126 

landslides along the Chengdu Baiyu section of the railway [2]. Landslide danger refers to the possibility of slope 

sliding and transforming into various forms of disasters under the combined action of multiple influencing factors 

[3]. The risk assessment of landslides can provide technical support for landslide prevention and control, and the 

reasonable selection of landslide evaluation factors is an important part of risk assessment.  

There have been studies on different methods for selecting factors and evaluating models in landslide risk 

assessment. Reference [4] used rough set, correlation analysis, and principal component analysis to screen 

landslide evaluation factors, and used support vector machine models for landslide susceptibility evaluation. 

Experimental examples showed that screening and reducing evaluation factors can improve the accuracy and 

accuracy of evaluation results; Literature [5] uses the Apriori algorithm to carry out correlation analysis on 

landslide evaluation factors and screen factors, and uses the Random forest model to evaluate landslide 

susceptibility. The experimental results show that the Apriori algorithm is used to select factors from the 

preselected factors that are more likely to cause landslides, and the evaluation results obtained are more consistent 

with the actual landslide distribution; Reference [6] used genetic algorithm and rough set to screen landslide 

evaluation factors, and used BP neural network for landslide susceptibility evaluation, indicating that the reduced 

factors correspond to higher accuracy of the evaluation model; BP neural network [7], support vector machine 

[8], Logistic regression [9] and Random forest [10] models have shown good performance in landslide 

susceptibility assessment; The literature [11] uses Random forest for landslide risk assessment, and compares it 

with Logistic regression, Multilayer perceptron and gradient enhancement tree. The results show that the accuracy 

rate of landslide risk assessment results obtained by Random forest model is the highest. In the above studies, in 

terms of selecting evaluation factors, some studies have directly selected factors related to landslide events for 

risk assessment, without considering the contribution of evaluation factors to landslide events, which can easily 

affect calculation efficiency and accuracy of evaluation results; In terms of evaluation factor screening, the data 

analysis method is often used to calculate the contribution rate of factors for factor screening. However, there are 

errors and Outlier in the evaluation factor data extracted from landslide events. When some important factors are 

eliminated due to data calculation and analysis, the evaluation results will be affected to a certain extent.  

To address the issue of mistakenly removing important evaluation factors based on their contribution rate, expert 

experience and calculation of evaluation factor contribution rate are introduced for factor screening. Firstly, the 

Analytic Hierarchy Process is used to calculate the subjective weights of the pre-selected factors. Then, grey 

correlation analysis is used to calculate the correlation degree of the pre-selected factors, and a distance function 

is introduced to calculate the comprehensive correlation degree of the evaluation factors based on subjective 

weights and grey correlation degree. Then, the evaluation factors are sorted according to the comprehensive 

correlation degree from large to small, and the first n factors are used as evaluation factor combinations, which 

mailto:topacademicjournals@gmail.com


    

 Academic Journal of Science, Engineering and Technology 

Vol.7, Issue 1; January - Febuary 2022; 

1252 Columbia Rd NW, Washington DC, United States 

https://topjournals.org/index.php/AJSET/index; mail: topacademicjournals@gmail.com 

  

 

 

3 | A c a d e m i c  J o u r n a l  o f  S c i e n c e ,  E n g i n e e r i n g  a n d  T e c h n o l o g y  

|  https://topjournals.org/index.php/AJSET 

are input into the Random forest model. The final evaluation factor combinations are selected according to the 

prediction accuracy of the model to complete the screening of evaluation factors, and the importance of evaluation 

factors is calculated through the Random forest. Finally, ArcGIS is used to stack the factor layers, calculate the 

landslide risk and generate the landslide risk distribution map in the study area, and complete the landslide risk 

assessment along the railway.  

2. Research methods  

2.1 Analytic hierarchy process  

The analytic hierarchy process (AHP) is a comprehensive evaluation model method. Through hierarchical and 

quantitative means, the relevant elements of decision-making problems are divided into multiple levels such as 

objectives, criteria, and indicators, and a multi-criteria analysis and decision-making method combining 

qualitative and quantitative analysis [12]. Its use steps are as follows:  

(1) The landslide risk assessment is selected as the target layer, the geological conditions, topographic 

conditions, and trigger conditions are selected as the criterion layer, and then the initial disaster causing factors 

are selected as the index layer to build the initial evaluation system;  

(2) The evaluation indexes in the same layer of landslide risk evaluation system are compared and scored, 

and the judgment matrix is constructed;  

(3) The weight w1 of each index is obtained by mathematical calculation;  

(4) The consistency index CI and random consistency index CR are introduced to test the consistency of each 

judgment matrix.  

λmax −n 

CI =                                    (1) n−1 

CI 

CR =                                      (2) RI 

2.2 Hierarchical grey relational analysis  

Landslide is a kind of disaster formed by the influence of many factors, in which the influencing factors are 

complex and vague, so the occurrence of landslide disaster can be regarded as a grey system. Grey correlation 

analysis measures the correlation degree between the comparison sample and the reference sample according to 

the similarity between the indicators [13], to determine the correlation degree of each indicator. The steps of grey 

correlation analysis are as follows:  

(1) Construct a sample matrix, a total of N groups of samples, each group of samples has a total of M 

indicators;  

(2) Determine the reference sample, which is composed of the optimal value or the worst value of each 

disaster causing factor;  

(3) Due to the different dimensions of a disaster causing factors, the normalized sample matrix is obtained by 

averaging them;  

mailto:topacademicjournals@gmail.com


    

 Academic Journal of Science, Engineering and Technology 

Vol.7, Issue 1; January - Febuary 2022; 

1252 Columbia Rd NW, Washington DC, United States 

https://topjournals.org/index.php/AJSET/index; mail: topacademicjournals@gmail.com 

  

 

 

4 | A c a d e m i c  J o u r n a l  o f  S c i e n c e ,  E n g i n e e r i n g  a n d  T e c h n o l o g y  

|  https://topjournals.org/index.php/AJSET 

 x0(1) 

 

x (2) 

（X X0, 1, Xn）=  0 

  

 

x m0(

 ) 

x1(1) 

x1(2) 

 

x m1(

 ) 

 

 

 

xn(1)  

 

xn(2)   

 

 x mn(

 )

 

(4) Calculate the absolute difference between each comparison sample and the reference sample, that is x k0( 

)− x ki ( ) ,k =1,2, ,m i; =1,2, ,n , secondly calculate the maximum and minimum  

values of all absolute differences;  

(5) Calculate the correlation coefficient;  

 n m n m 

minmin x k0 ( )− x ki +ρmaxmax x k0 ( )− x ki 

 ξi ( )k = i=1 k=1 n i=m1 k=1                (3) x k0 ( )− x ki +ρmaxmax x k0 ( )− x ki 

 i=1 k=1 

In equation (3) p is the resolution coefficient, generally 0.5[14].  

(6) The average method is used to calculate the correlation degree;  

1 m 

ri 
= ∑ξi ( )k                               (4) m k=1 

(7) The normalized grey correlation degree W2 is obtained by normalizing the correlation degree;  

(8) The distance function[15] is introduced to calculate the comprehensive correlation degree combined with 

AHP weight and normalized grey correlation degree.  

，w2 ) = 1 ∑m (wi1 −wi2 )2                         (5)  

L w( 1 

2 i=1 

Let the comprehensive correlation degree be w and the coefficients of AHP weight and grey correlation degree be 

a and b respectively, then w is:  

w = aw1 +bw2                                (6)  

Make the difference between grey correlation degree and AHP weight consistent with the difference between 

distribution coefficients, to eliminate the influence of human factors and data singularity, and obtain relatively 

objective and accurate results. Therefore, the following constraints are used to calculate the correlation coefficient.  

 L w2 ( 1，w2 ) = −(a b)2 

                              (7)  

  a+ =b 1 

mailto:topacademicjournals@gmail.com


    

 Academic Journal of Science, Engineering and Technology 

Vol.7, Issue 1; January - Febuary 2022; 

1252 Columbia Rd NW, Washington DC, United States 

https://topjournals.org/index.php/AJSET/index; mail: topacademicjournals@gmail.com 

  

 

 

5 | A c a d e m i c  J o u r n a l  o f  S c i e n c e ,  E n g i n e e r i n g  a n d  T e c h n o l o g y  

|  https://topjournals.org/index.php/AJSET 

2.3 Association rules  

Association rules are one of the important means of data mining, which mainly refers to finding the correlation 

of different items in the same event, that is, the correlation between different items in the transaction database 

[16]. The total set of the chronicle database is X. if subsets A and B of the total set X meet A⊂X,B⊂X, and 

A∩B=⌀, A⇒B is called association rule, where a and B are the premise and consequence of association rule 

respectively. Support is the percentage of A∪B in transaction database x, as shown in equation (8):  

 Support A( => =B) p A( B)                         (8)  

Calculate the support degree of landslide caused by each disaster causing factor, and use the support degree to 

test the rationality and accuracy of landslide disaster causing factor screening [17].  

2.4 Random forest  

Random forest (RF) is a combined classifier in machine learning algorithm. N samples (n < N) are retrieved from 

n sample sets, K attributes (k < K) are selected from K total attributes, and the best segmentation attributes are 

selected based on Gini index to create a decision tree. Multiple decision trees can be integrated through bagging 

algorithm to form a random forest [18]. Random forest training is fast, not easy to "over fit", and has a good 

tolerance for noise and outliers [19]. The importance of RF is calculated based on gini index. The calculation 

formula of gini index is as follows:  

n 

gini A( ) = −1 ∑ pi
2                                 (9)  

i=1 

m 

gini A B( , ) =∑gini A( j )                            (10)  

j=1 

In formula (9), n is the number of classification categories in sample A (whether landslide disaster occurs); pi is 

the proportion of class i samples in sample A. In equation (10), gini (A, B) is the gini index after dividing sample 

A by factor B; m is the total number of samples| Aj | is the number of j samples. After averaging the importance 

results output from each decision tree, it is the importance of the random forest.  

3. Landslide risk assessment model  

The flow chart of landslide risk assessment is shown in Figure 1.  

j A 

A 

mailto:topacademicjournals@gmail.com


    

 Academic Journal of Science, Engineering and Technology 

Vol.7, Issue 1; January - Febuary 2022; 

1252 Columbia Rd NW, Washington DC, United States 

https://topjournals.org/index.php/AJSET/index; mail: topacademicjournals@gmail.com 

  

 

 

6 | A c a d e m i c  J o u r n a l  o f  S c i e n c e ,  E n g i n e e r i n g  a n d  T e c h n o l o g y  

|  https://topjournals.org/index.php/AJSET 

 
  

Figure 1: Flow chart of landslide risk assessment  

(1) Firstly, score the initial evaluation index system and construct the judgment matrix, so as to calculate the 

subjective weight of each factor and test the consistency of the judgment matrix. At the same time, the grey 

correlation method is used to analyze the initial indexes and calculate the grey correlation degree of each factor. 

The comprehensive correlation degree is calculated by using the distance function,  

(2)The factors are sorted according to the comprehensive correlation degree from large to small. Take the first n 

factors and input them into the random forest model. Select the corresponding factor combination with the highest 

accuracy, eliminate the factors that do not appear in the factor combination, and complete the factor screening. 

Then the association rules are used to calculate the support of each factor to verify the rationality and accuracy of 

factor screening.  

(3) After the factor screening, use the random forest to calculate the importance of each factor, set and optimize 

multiple parameters in the random forest, improve the classification accuracy of the model, and obtain the 

importance of each factor. Use ArcGIS to assign values and overlay analysis to each factor layer, calculate the 

landslide risk size and generate the landslide risk distribution map, and complete the landslide in Ya'an-Batang 

section of Sichuan-Tibet railway Risk assessment.  

4. Research area and data source  

4.1 Overview of the study area  

Figure 2 is the railway route map of Ya'an City. Ya'an City is complex and changeable, the fault structure is 

developed, the geological disasters are many and wide, the dangerous situation is serious and the harm is great, 

and the rainfall frequency is intensive from June to August every year. The rainfall is mostly heavy rain and heavy 

rainstorm, and the landslide types are mostly rainfall type landslide and earthquake type landslide. The triggering 

Analytic  
hierarchy process 

Grey relation  
analysis 

Comprehensive  
relevance ranking 

Initial landslide risk  
evaluation index system 

Random forest 

Landslide risk  
assessment 

mailto:topacademicjournals@gmail.com


    

 Academic Journal of Science, Engineering and Technology 

Vol.7, Issue 1; January - Febuary 2022; 

1252 Columbia Rd NW, Washington DC, United States 

https://topjournals.org/index.php/AJSET/index; mail: topacademicjournals@gmail.com 

  

 

 

7 | A c a d e m i c  J o u r n a l  o f  S c i e n c e ,  E n g i n e e r i n g  a n d  T e c h n o l o g y  

|  https://topjournals.org/index.php/AJSET 

conditions of landslide disaster in this study only consider rainfall and human engineering activities, not the 

earthquake.  

  
Figure 2: Study area  

4.2 Data sources  

Based on the existing research[20,21], 12 disaster causing factors are selected, namely rainfall, formation 

lithology, distance to fault, land use type, slope, elevation, slope direction, plane curvature, profile curvature, 

distance to river, vegetation coverage and human engineering activities. These factors can comprehensively 

present the geological conditions, topographic conditions, and triggering conditions of landslide disasters in the 

study area.  

The data sources are as follows: the rainfall data is from the China Meteorological data network. The maximum 

24h rainfall in the current rainfall process is selected to represent the rainfall factor; The data of formation 

lithology and distance to fault are from the national geological data center; Land use types and landslide disaster 

distribution are derived from the resource and environmental science and data center; The distance to the river is 

derived from the basic geographic information database; Human engineering activities are represented by road 

buffer zone and residential area, and the road buffer zone originates from OSM(https://www.openstreetmap.org/), 

the residential area originates from network collection; Landsat8 and DEM data are derived from geospatial data 

cloud, vegetation coverage is extracted from landsat8, and slope, aspect, plane curvature, and section curvature 

are extracted from  

DEM.  

5. Example analysis and discussion  

5.1 Subjective weight calculation based on AHP  

The initial landslide risk evaluation index system is constructed according to 12 initial hazard factors, as shown 

in Figure 3. The evaluation indexes are scored, the judgment matrix A is constructed based on the geological 

conditions, topographic conditions and trigger conditions in Figure 3, and the judgment matrices B1, B2 and B3 

mailto:topacademicjournals@gmail.com


    

 Academic Journal of Science, Engineering and Technology 

Vol.7, Issue 1; January - Febuary 2022; 

1252 Columbia Rd NW, Washington DC, United States 

https://topjournals.org/index.php/AJSET/index; mail: topacademicjournals@gmail.com 

  

 

 

8 | A c a d e m i c  J o u r n a l  o f  S c i e n c e ,  E n g i n e e r i n g  a n d  T e c h n o l o g y  

|  https://topjournals.org/index.php/AJSET 

are constructed based on the hazard factors under geological conditions, topographic conditions and trigger 

conditions.  

 
  

Figure 3: Initial landslide risk evaluation index system  

6 3 4 4 4 4 4 2 1 

3 2 1 5 2 3 5 3 1 5 3 3 

4 3 5 1 11 9 3 4 1 2 

 4 5     5 1 5 2   6 10 10 4 3   B3= 1   

 A= 1 4 3 2 10 4 2 9 1 

 3 4 B1= 1 2 

 3 3 6 3 11 5 5 10 

3 4 1 B2= 

1 4 5 5 3 10 5 2 6 

2 
5 3 1 5 4 9 4 1 3 5 

1 

 2 2 6 5 4 5 3 1 2 

4 3 2 2 

5 3 10 5 1 

1 

7 4 9 6 2 

The eigenvectors of each judgment matrix are calculated, and the eigenvectors of each disaster causing factor are 

obtained importance w1, check the consistency of each judgment matrix through equations (1) and (2), as shown 

in Table 1.  

mailto:topacademicjournals@gmail.com


    

 Academic Journal of Science, Engineering and Technology 

Vol.7, Issue 1; January - Febuary 2022; 

1252 Columbia Rd NW, Washington DC, United States 

https://topjournals.org/index.php/AJSET/index; mail: topacademicjournals@gmail.com 

  

 

 

9 | A c a d e m i c  J o u r n a l  o f  S c i e n c e ,  E n g i n e e r i n g  a n d  T e c h n o l o g y  

|  https://topjournals.org/index.php/AJSET 

w1=(0.062,0.081,0.064,0.054,0.057,0.055,0.049,0.058,0.87,0.052,0.253,0.127)  

Table 1: Consistency test results of each judgment matrix  

  A  B1  B2  B3  

λ  

CI  

CR 

3.0015 

0.0007  

0.0015  

4.0471 

0.0157  

0.0176  

6.0300 

0.0060  

0.0047  

2  

0  

0  

5.2 Calculation of correlation degree of disaster causing factors based on grey relation  

12 preselected factors such as fault distance and formation lithology are selected as the indexes of grey correlation 

analysis. The extracted initial landslide data are processed by the mean method, and the reference samples are 

determined:  

X0=(1.6,2,1.02,1.33,1.42,1.6,2.3,1.2,1.1,1.2,1.42,1.17)  

Calculate the absolute difference between the comparison sample and the reference sample, and determine the 

maximum and minimum values of the absolute difference. The correlation degree is calculated according to 

equations (3) and (4):  

r=(0.872,0.548,0.499,0.476,0.831,0.85,0.502,0.899,0.923,0.542,0.893,0.881) The grey relation degree is 

normalized to:  

r1=(0.1,0.063,0.057,0.055,0.095,0.098,0.058,0.103,0.106,0.062,0.102,0.1) The comprehensive relation degree 

obtained from equation (5) is:  

Z=(0.081,0.071,0.06,0.054,0.076,0.076,0.053,0.08,0.096,0.055,0.177,0.11)  

5.3 factor screening and inspection  

Rank the factors in the comprehensive correlation degree z from large to small, Take the first n factors as data 

sets and input them into the random forest model (n = 5, 6,..., 12). ROC curve is called sensitivity curve, which 

has been widely used in the accuracy analysis of geological hazard risk assessment results[22]。 AUC represents 

the area under the ROC curve. The closer the AUC value is to 1, it shows that the discrimination result of the 

model is better. As shown in Figure 4, after inputting the random forest model through different factor 

combinations, the ROC curve is generated and the AUC value is calculated. When n = 8, the corresponding model 

accuracy is the largest, and the AUC value is 0.89. Therefore, keep these 8 factors, eliminate the distance to fault, 

distance to the river, plane curvature, and section curvature, and complete the factor screening. The association 

rule formula (8) is used to calculate the support of the above-retained factors, and the obtained support is 0.94, 

0.941, 0.824, 0.912, 0.81, 0.891, 0.906, and 0.912, which are greater than 0.8. Therefore, this method is reasonable 

for screening landslide risk assessment indicators.  

mailto:topacademicjournals@gmail.com


    

 Academic Journal of Science, Engineering and Technology 

Vol.7, Issue 1; January - Febuary 2022; 

1252 Columbia Rd NW, Washington DC, United States 

https://topjournals.org/index.php/AJSET/index; mail: topacademicjournals@gmail.com 

  

 

 

10 | A c a d e m i c  J o u r n a l  o f  S c i e n c e ,  E n g i n e e r i n g  a n d  T e c h n o l o g y  

|  https://topjournals.org/index.php/AJSET 

  
Figure 4: Test results of random forest ROC curve of eight factor combinations  

Python is used to calculate the importance of each factor in the random forest model. Firstly, the size of each 

parameter is set, and then the parameters are adjusted. The optimal values of each parameter are obtained by grid 

search and 10 fold cross verification. Finally, the importance w2 of the eight factors of rainfall, human engineering 

activities, slope, stratum lithology, vegetation coverage, slope direction, elevation, and land use type is obtained 

through feature_importances function.  

w2=(0.201,0.17,0.154,0.141,0.102,0.096,0.072,0.063)  

5.4 landslide risk assessment in the study area  

On the basis of each factor layer, the important weight is given to each layer by using the spatial superposition 

function of ArcGIS, and the landslide risk distribution map of the study area is generated, as shown in Figure 5. 

Based on the natural breakpoint method, the landslide risk evaluation results are divided into five levels: low, low, 

medium, high and high risk. It is generally believed that the high-risk and high-risk areas are vulnerable to 

landslides. When 97 landslide verification points are used to test the landslide risk assessment results, it can be 

obtained that 89.7% of the landslide points are located in the high-risk and high-risk areas, that is, it is considered 

that the verification accuracy is 89.7%. According to the statistical analysis of the risk assessment results in Figure 

5, the railway lines in high and high-risk areas are 10.9 km (2.1%) and 48.7 km (9.4%) respectively and are mostly 

distributed in Ya'an area. The main reason is that Ya'an has large rainfall, which is easy to reduce the shear strength 

of the slope and accelerate the disintegration and destruction of the slope. Among them, the maximum 24h rainfall 

in August 2020 is 425.2mm, Reaching an all-time high. At the same time, human engineering activities in this 

area are strong, roads and residential areas are densely distributed, which provides a good disaster pregnant 

environment for the landslide. For the above high-risk lines, it is recommended to use scientific means to monitor 

the dangerous slopes during line construction, and take preventive measures against landslides in advance, so as 

to prevent landslides from causing huge damage to the railway line. During the later operation of Sichuan-Tibet 

railway, the landslide hazard assessment results are closely combined with the railway signal early warning system 

to ensure the safe operation of trains and the safety of lines, and provide a safety guarantee for railway operation.  

mailto:topacademicjournals@gmail.com


    

 Academic Journal of Science, Engineering and Technology 

Vol.7, Issue 1; January - Febuary 2022; 

1252 Columbia Rd NW, Washington DC, United States 

https://topjournals.org/index.php/AJSET/index; mail: topacademicjournals@gmail.com 

  

 

 

11 | A c a d e m i c  J o u r n a l  o f  S c i e n c e ,  E n g i n e e r i n g  a n d  T e c h n o l o g y  

|  https://topjournals.org/index.php/AJSET 

  
Figure 5: Landslide risk zoning map of the study area  

5.5 Accuracy verification of different models  

The random forest model is combined with the support vector machine (SVM) model and logistic regression 

model commonly used in landslide risk assessment (LR) to compare the model accuracy before and after 

screening factors. As shown in Figure 6, the accuracy of random forest model before screening factors is the 

largest, and its AUC area is 0.82. After comprehensive correlation screening factors and random forest parameter 

optimization, the AUC area of random forest model is increased to 0.91, which is greater than 0.81 of logistic 

regression and 0.85 of support vector machine.  

   
(a)Test results of ROC curve of each model without screening factor  

mailto:topacademicjournals@gmail.com


    

 Academic Journal of Science, Engineering and Technology 

Vol.7, Issue 1; January - Febuary 2022; 

1252 Columbia Rd NW, Washington DC, United States 

https://topjournals.org/index.php/AJSET/index; mail: topacademicjournals@gmail.com 

  

 

 

12 | A c a d e m i c  J o u r n a l  o f  S c i e n c e ,  E n g i n e e r i n g  a n d  T e c h n o l o g y  

|  https://topjournals.org/index.php/AJSET 

  
(b) ROC curve test results of each model after screening factors  

Figure 6: ROC curve verification  

6. Conclusion  

Based on the grey correlation and AHP screening factors, the random forest model is used to evaluate the landslide 

risk. Taking Ya'an city railway lines as the research object, the following conclusions are obtained:  

(1) The grey relation and analytic hierarchy process are used to comprehensively screen the factors, which 

solves the problem of false deletion of factors due to a small grey correlation degree. The results show that rainfall, 

slope, human engineering activities, formation lithology, vegetation coverage, slope direction, elevation, and land 

use type are strongly related to the occurrence of landslide disaster in Ya'an city railway lines.  

(2) The random forest model is used to evaluate the landslide risk, and the ROC curve is used to evaluate the 

accuracy of the random forest model. The area under the ROC curve before the screening factor is 0.82, and the 

area under the ROC curve after the optimization of the screening factor and random forest parameters is 0.91. It 

is proved that the optimization of the screening factor and model parameters can improve the accuracy of the 

evaluation model.  

(3) The landslide risk assessment results obtained based on the research method in this paper are consistent 

with the landslide disaster verification points in Sichuan Province, which proves that the research method in this 

paper is reliable, accurate and has certain practical value. It can provide some result reference and technical 

support for the medium-term construction and later operation management of Ya'an Railway.  

References  

Lu CF, Cai CX, 2019. Challenges and Countermeasures for Construction Safety during the Sichuan–Tibet 

Railway Project. Engineering, 5(05):49-61. DOI: CNKI:SUN:GOCH. 0. 2019-05-010.  [2] Li XZ, Cui Y, 

Zhang XG, et al., 2019. Characteristics and spatial distribution law of landslides and collapses along 

Sichuan-Tibet railway. Journal of Engineering Geology, 27(S): 110-120. DOI:10. 26914/c. cnkihy. 2019. 

014704.   

mailto:topacademicjournals@gmail.com


    

 Academic Journal of Science, Engineering and Technology 

Vol.7, Issue 1; January - Febuary 2022; 

1252 Columbia Rd NW, Washington DC, United States 

https://topjournals.org/index.php/AJSET/index; mail: topacademicjournals@gmail.com 

  

 

 

13 | A c a d e m i c  J o u r n a l  o f  S c i e n c e ,  E n g i n e e r i n g  a n d  T e c h n o l o g y  

|  https://topjournals.org/index.php/AJSET 

Wang Ting, 2020. Landslide Risk Assessment Based on Combined Analysis of Weight Method and Simulation 

Method: A Case Study of Parts of Wanyuan City. DOI: 10. 26986/d. cnki. gcdlc. 2020. 000805.   

Yu XY, Hu YJ, Niu RQ, 2016. Research on the Method to Select Landslide Susceptibility Evaluation Factors Based 

on RS-SVM Model. Geography and Geo-Information Science, 32(03): 23-28+2. Doi: CNKI: SUN: DLGT. 

0. 2016-03-005.   

Liu R, Li LY, Yang X, et al., 2021, Landslide Susceptibility Mappping Based on Factor Strong Correlation Analysis 

Method[J]. Earth and Environment, 49(02):198-206. Doi: 10. 14050/j. cnki. 1672-9250. 2020. 48. 108.   

Tang RX, Yan EC, Tang W, 2017. Landslide susceptibility evaluation based on rough set and back-propagation 

neural network [J]. Coal Geology & Exploration, 45(06):129-138.   

Aditian A, Kubota T, Shinohara Y, 2018. Comparison of GIS-based landslide susceptibility models using 

frequency ratio, logistic regression, and artificial neural network in a tertiary region of Ambon, Indonesia. 

Geomorphology, 318:101-111. DOI:10. 1016/j. geomorph. 2018. 06. 006.   

Huang Y, Zhao L. 2018. Review on landslide susceptibility mapping using support vector machines. Catena, 165: 

520−529. DOI:10. 1016/j. catena. 2018. 03. 003.   

Luigi L, Martin M P. 2018. Presenting logistic regression-based landslide susceptibility results. Engineering 

Geology, 244: 14−24. DOI:10. 1016/j. enggeo. 2018. 07. 019.   

Wu XQ, Lai CG, Chen XH, et al. A landslide hazard assessment based on random forest weight:a case study in 

the Dongjiang River Basin, Journal of Natural Disasters, 2017, 26(05): 119-129. DOI: 10. 13577/j. jnd. 

2017. 0514.   

Ahmed N, Firoze A, Rahman RM. 2020. Machine learning for predicting landslide risk of Rohingya refugee camp 

infrastructure. Journal of Information and Telecommunication, 4(2): 175-198. DOI: 10. 1080/ 24751839. 

2019. 1704114.   

Wang JH, Jin HL, Ni TX, et al., 2017. The application of fuzzy comprehensive evaluation model based on analytic 

hierarchy process in risk assessment of debris flow gully in Kangle County. The Chinese Journal of 

Geological Hazard and Control, 28(03):52-57. DOI: 10. 16031/j. cnki. issn. 1003- 8035. 2017. 03. 08.   

Huang F, Wang Y, Dong Zhiliang, et al., 2019. Regional Landslide Susceptibility Mapping Based on Grey 

Relational Degree Model. Earth Science, 44(02):664-676.   

mailto:topacademicjournals@gmail.com


    

 Academic Journal of Science, Engineering and Technology 

Vol.7, Issue 1; January - Febuary 2022; 

1252 Columbia Rd NW, Washington DC, United States 

https://topjournals.org/index.php/AJSET/index; mail: topacademicjournals@gmail.com 

  

 

 

14 | A c a d e m i c  J o u r n a l  o f  S c i e n c e ,  E n g i n e e r i n g  a n d  T e c h n o l o g y  

|  https://topjournals.org/index.php/AJSET 

Jaiswal P, Van. W C J, Jetten V. Quantitative assessment of direct and indirect landslide risk along transportation 

lines in southern India[J]. Natural hazards and earth system sciences, 2010, 10(6):1253-

1267.DOI:10.5194/nhess-10-1253-2010.  

Zhang C, Wang Q, Chen JP, et al., 2011. Evaluation of debris flow risk in Jinsha River based on combined weight 

process. Rock and Soil Mechanics, 32(03):831-836. DOI: 10. 162 85/ j. rsm. 2011. 03. 019.   

Tan GS, Cao SX, Zhao B, et al., 2020. An assessment of power transformers based on association rules and 

variable weight coefficients. Power System Protection and Control, 48(01): 88-95. DOI: 10. 19783/j. cnki. 

pspc. 190072.   

Song RJ, Liu RY, Liu YW, et al., 2018. Key Indicator System Establishment and Application for Transformer 

Condition Assessment Based on Hierarchy Grey Relation Analysis. High Voltage Engineering, 

44(08):2509-2515. DOI: 10. 133 36/j. 1003-6520. hve. 20180731010.   

Zhang WQ, Luo GP, Zheng HW, et al., 2020. Analysis of vegetation index changes and driving forces in inland 

arid areas based on random forest model: a case study of the middle part of northern slope of the north 

Tianshan Mountains. Chinese Journal of Plant Ecology, 44(11): 1113-1126.   

Breiman L. 2001. Random forests. Machine Learning, 2001, 45(1): 5-32. DOI: 10. 1023/A: 1010933404324.   

Yang ZJ, Ding PP, Wang D, et al., 2018. Landslide Risk Analysis on et a Sichuan-Tibet Railway (Kangding to 

Nyingchi Section) Journal of the China Railway Society, 40(09): 97-103. DOI: CNKI: SUN: TDXB. 0. 

2018-09-015.   

Wang WD, Fu QX, Tang R. 2018. Landslide spatiotemporal susceptibility analysis of Chengdu-Yaan section in 

Sichuan-Tibet railway. Journal of Railway Science and Engineering, 15(04): 862-870. DOI:10. 19713/j. 

cnki. 43-1423/u. 2018. 04. 006.   

Reichenbach P, Rossi M, Malamud B D, et al. A review of statistically-based landslide susceptibility models [J]. 

Earth-Science Reviews, 2018, 180, 60-91.   

mailto:topacademicjournals@gmail.com

