Frontiers in Business, Economics and Management ISSN: 2766-824X | Vol. 6, No. 1, 2022 177 Research on Network Customer Classification Based on BP Neural Network Algorithm Hankun Ye School of International Trade and Economics, Jiangxi University of Finance and Economics, Nanchang, 330013, China Abstract: Correctly and effectively customer classification according to their characteristics and behaviors will be the most important resource for electronic marketing and online trading of network enterprises. A new customer classification model for online trading customer classification is presented based on analyzing customer characteristics and behaviors .First the paper designs 21 customer classification indicators based on consumer characteristics and behaviors analysis, including customer characteristics type variables and customer behaviors type variables; Second, Aiming at the shortages of the existing BP neural network algorithm of data-mining for customer classification, immune genetic algorithm is used to correct BP neural network to speed up the convergence of the model. Finally the experimental results verify that the new algorithm can improve effectiveness and validity of customer classification when used for classifying network trading customers practically. Keywords: Electronic marketing, Customer classification, BP neural network algorithm, Immune Genetic Algorithm. 1. Introduction Customer relations management is one of the core problems of modern enterprises, whose customer oriented thought requires CRM system to be able to effectively obtain various kinds of information of customers, identify all the relations between the customers and enterprises and understand the transaction relation between customers and enterprises; meanwhile, deeply analyze customers’ consuming behavior, find customers’ consumption characteristics, providing personalized service for customers, supporting the decisions of enterprises. The three basic problems CRM needs to solve are how to get customers, how to keep customers and how to maximize customer value, among which maximizing customer value is the ultimate purpose, getting customers and keeping customers are both the means for realizing the purpose. The core of analyzing the three problems CRM needs to solve is to classify customers. “Getting Customers” and “Retaining Customers” need to ascertain which customers are attainable, which customers need to be kept, which customers are kept for a long term and which customers are kept for a short term, therefore, customer classification is needed. It is the same case with “Maximizing Customer Value”. Due to different values of different customers, “Maximum Customer Value” of different customers should be distinguished. Thus, the core problem of enterprises to correctly implement CRM is to adopt effective method to reasonably classify customers, find customer value, focus on high-value customers with enterprises’ limited resources, provide better service for them, keep “High-value” customers for loss prevention; also, establish corresponding customer service system through classification, carry out differential customer service management. Hence, customer classification is becoming a more and more popular research hotspot, also a research difficulty, becoming one of the urgent problems of CRM[1]. 2. Selection of Customer Classification Indicators The selection of reasonable classification variables is the basis of correct and effective customer classification, namely establishing scientific and reasonable classification indicators system. In view of the nature of trading and own characteristics of online trading, this Paper adopts customer characteristics type variable and customer behaviors type variable in the specific selection of customer classification variables[2]. (1) Selection of Customer Characteristics Type Variable Customer characteristics type variable is mainly used for getting the information of customers’ basic attributes. Such variable indicators as geographical position, age, sex, income of individual customer play a key role in determining the members of some market segment. This kind of variables mainly comes from customers’ registration information and customers’ basic information collected from the management system of banks, the contents of which mostly indicate the static data of customers’ basic attributes, the advantage of which is that most of the contents of variables are easy to collect. But some of the basic customer-described contents of variables are lack of differences at times. Based on analyzing and summarizing existing literatures, the customer characteristics type variables designed in this Paper include: Customer No., Post Code, Date of Birth, Sex, Educational Background, Occupation, Monthly Income, Time of First Website Browsing, and Marital Status. (2) Selection of Customer Behaviors Type Variables Customer behaviors type variables mainly indicate a series of variable indicators related to customer transacting behavior and relation with banks, which are used to define the orientation which enterprises should strive for in some market segment, and are the key factors for ascertaining target market. Customer behaviors type variables include the records of customers buying services or products, records of customer service or production consumption, contact records between customers and enterprises, as well as customers’ consuming behaviors, preferences, life style, and other relevant information. Based on analyzing and summarizing existing literatures, the customer behaviors type variables designed in this paper include Monthly Frequency of Website Login, Monthly Website Staying Time, Monthly Times of Purchasing, 178 Monthly Amount of Purchasing, Type of Consumer Products Purchased, Times of Service Feedback, Service Satisfaction, Customer Profitability, Customer Profit, Repeat Purchases, Recommended Number of Customers, Purchasing Growth Rate. 3. Derivation of Algorithm 3.1. Simultaneous Analysis and Design De Castro indicated that there were similarities among the quality of weight value initialization of back-propagation neutral network and the relationship of network output and the quality of antibody instruction system initialization in the immune system and the quality of immune response. A simultaneous analysis and design---SAND algorithm was advanced to solve the problem regarding the weight value initialization in the back-propagation network[7,8]. In SAND algorithm, each antibody corresponds to a weight value vector of neuron given in one of several layers of neural networks, the length is l , and the affinity ),( ji xxaff between antibody ix and antibody jx is shown by their derivative of Euclidean distance function ),( ji xxD in Formula 1. In which,  is a positive of value adoption 0.001. The definition of Euclidean distance function ),( ji xxD is shown in Formula 2[8].   ),( 1 ),( ji ji xxD xxaff (1)    l k jkikji xxxxD 1 2)(),( = (2) SAND algorithm aims to reduce the similarities between the antibodies and produce the antibody repertoire to cover the entire form space with the best, so energy function is maximized. The energy function is shown in Formula 3.      N i N ij ji xxDE 1 1 ),( (3) In the method of Eculidean form space, the energy function is not percentage. With a view to the diversity of the vector, SAND algorithm has to define the stop condition. Given vector Nixi ,...2,1,  , its standardization is unit vector NiI i ,...2,1,  ,  I shows to calculate the average vector. Therefore, Formula 4 shows the diversity of unit vector, in which,  I means the average vector distance from the origin of coordinate. Formula 5 shows the stop conditionU of SAND algorithm. 2/1)( III T  (4) )1(100   IU (5) 3.2. BP Neural Network Design Based on Immune Genetic Algorithm According to the actual application, providing that both the input and output number of node and the input and output values in BPNN have been confirmed, activation function adopts S type function. The following steps show BP neural network design based on immune genetic algorithm. (1) Every layer of BPNN carries on the weight value initialization separately by SAND algorithm. (2) Antibody code. The initial weight value derived by SAND algorithm constructs the structures of BPNN. Each antibody corresponds to a structure of BP neural network. The number of hidden node and network weight value carry on the mixture of real code. Each antibody serials are shown in Fig.3. N number of hidden node Weight value corresponding to the first hidden node Weight value corresponding to the second hidden node … Weight value corresponding to the N hidden node Figure 3. Antibody Code (3) Fitness function design. Fitness function )( ixf is defined as the mean value function of squared error of neural network in Formula 6, in which, )( ixE is shown by Formula 7. In Formula 7, p is the total training sample, o is the number of node of output layer, n jT and n jY are the n training sample’s expected output and actual output in the j output node separately, and  is the constant larger than zero. )( 1 =)( i i xE xf (6) 2 1 1 )( 2 1 )( nj p N o j n ji YT p xE      (7) (4) Genetic operation. The model here adopts the Gaussian compiling method to go on the genetic operation so as that each antibody decoding is the corresponding network structure and change the network weight value as shown in Formula 8, in which, ix and m ix are the antibodies before and after the variation, ),( 10 shows that the mean value is zero and squared error is normal distribution random variable of l , and )1,1( is the individual variation rate. It is seen in Formula 8 that the variation degree varies inversely as the fitness, i.e. the lower the fitness is (the less 179 the fitness value of objective function is), the higher the individual variation rate is, or vice versa. After the variation, all the hidden node and weight value components constitute a new antibody again. )1,0())(exp(  ii m i xfxx (8) (5) Group renewal based on density. In order to guarantee the antibody diversity, improve the entire searching ability of the algorithm, the model adopts the Euclidean distance and the fitness based on the antibodies to calculate the similarity and density of the antibody. Providing that there are ix and jx antibodies, and 0 and 0t given constants, the fact that Formula 9 is satisfied indicates that ix and jx antibodies are similar, the number of antibody similar to the antibody ix is the density of ix marked by iC . The probability of selecting antibody ix is )( ixp as shown in Formula 10, in which,  and  is the adjustable parameters between (0, 1), and )( xM is the maximum fitness value of all the antibodies. It is seen in Formula 10 that while the antibody density is high, the probability of selecting the antibody with high fitness is low, and conversely high. Therefore, excellent individual is not only retained, but the selection of similar antibodies is reduced, and the individual diversity is guaranteed.       txfxf xxD ji ji )()( ,(  (9) )( )( ] )( )( 1[)( xM xf xM xf Cxp iiii   (10) 4. Experimental Verification 4.1. Object of Experimental Verification The instance data of the experiment conduct empirical research on the customer data of the B2C transaction of certain enterprise website of the recent three years (totaling data of 41351 customers, 21 attributes in the data table are listed in the third part of the paper including customer characteristics type variables and customer behaviors type variables), making statistics on attribute values like annual transaction frequency, total amount, product cost, etc. of certain customer according to customer transaction records in information base, forming an information table (among which the decision attribute set D is null) [4]. 4.2. Process of Experimental Verification The process of the experimental verification can be listed as follows[5]. First, what is to be processed during the classification is the numeric data, so the numeric coding on character data should be conducted first; Second, if the value number of certain attribute is equal to sample number, it means that it has little effect on classification, hence, remove such attribute first. Three attributes as Customer No., Post Code and Date of Birth are removed in this case. Third, establish training sample set according to domain (prior) knowledge. Times of purchasing and total amount of purchasing of each customer are two major factors of customer classification (this is the prior knowledge of domain), so select 400 pieces of typical data among all the customers to form training sample set. And divide them into five types as Gold Customers, Silver Customers, Copper Customers, General Customers and Negligible Customers according to ABC management theory. Fourth, use the customer classification algorithm above- mentioned, and the customer classification results can be expressed in Table 1. In the specific algorithm realization, this Paper simultaneously realizes ordinary K_means algorithm and customer classification algorithm based on BP neural network. The performance comparison of these three algorithms can be expressed in Table 2. Table 1. Customer Classification Result of Some Website Customer Type Number of Customers Percentage % Profit Contribution Proportion Gold Customers 2849 6.89 52.1 Silver Customers 5921 14.32 30.1 Copper Customers 10193 23.65 13.1 General Customers 13751 36.44 6.1 Negligible Customers 7319 17.70 -1.4 Total 41351 100.00 100.00 We can see from Table 1 that in the autonomous learning of algorithm of this Paper, such five factors as the educational background, income, occupation, times of purchasing, and total amount of purchasing of customers have a relatively great influence on customer classification. Through the classification result in Table 1, it can be seen that Gold Customers take up 6.89% of the total number of customers, while the profit takes up 52.1% of the total profit. These customers play a significant role in the existence and development of enterprises. However, the negligible customers account for 17.7%, who not only do not bring profit to enterprises, but also make enterprise lose money. These customers should be either further cultivated or eliminated according to the actual situation. 180 Table 2. Classification Performance Comparison of Each Algorithm Algorithm Algorithm in This Paper Ordinary K-means Algorithm BP Neural Network Algorithm Accuracy Rate 99.7 % 88.47% 94% E Value 104.33 159.81 119.96 We can see from Table 2 that the cluster accuracy rate of algorithm in this paper is the highest, reaching 99.7 %, obviously higher than ordinary K-means algorithm and BP Neural Network algorithm; the square errors and E values on customer classification of three algorithms are 104.33, 159.81 and 119.96 respectively. The smaller the E value is, the smaller the possibility of wrong classification is. Thus it can be seen that the square error and E value of the algorithm in this paper during the classification are far more less than ordinary K-means algorithm[4] and BP Neural Network algorithm[6]. Therefore, it shows that the improvement on K-means clustering algorithm in this paper turns out to be a success, with reasonable classification results. 5. Conclusion Customer relations management of online trading is still developing. But to correctly and effectively classify online trading customers is the critical issue for reforming network marketing mode, improving customer management and service level and enhancing competitiveness of network enterprise. On account of the shortcomings of the typical K_means clustering algorithm in data mining, this Paper puts forward several improvement measures, and applies them into the classification of online trading customers. Simulation results indicate that the improved online trading customer classification has higher accuracy rate on customer classification and more reasonable classification results. References [1] Liu Zhaohua. Study on Model of Customer Classification Based on the Customer Value, A Dissertation of Huazhong University of Science and Technology, 2021. [2] Zhou Huan, Study of Classifying Customers Method in CRM, Computer Engineering and Design, 2018,Vol 29, No.3, pp.659- 661. [3] Deng Weibing, Wang Yan, B2C Customer Classification Algorithm Based on Based on 3DM, Journal of Chongqing University of Posts and Telecommunications (Natural Science Edition), 2019,Vol 21, No.4, pp.568-572. [4] Guan Yunhong, Application of Improved K-means Algorithm in Telecom Customer Segmentation, Computer Simulation, 2022, Vol 28, No.8, pp.138-140. [5] Qu Xiaoning, Application of K-means Based on Commercial Bank Customer Subdivision, Computer Simulation, 2018, Vol 28, No.6, pp.357-360. [6] Yang Benzhao, Tian Gen, Research on Customer Value Classification based on BP Neural Network Algorithm, Science and Technology Management Research, 2017, Vol 23, No.12, pp.168-170. [7] ,Bradley P S Managasarian .L k- .plane Clustering Journal ,of Global Optimization 202 ,0 16(1)23-32. [8] ,Tang Yong Rong Qiusheng. An Implementation of Clustering Algorithm Based on K-means. Journal of Hubei Institute for Nationalities, 2014, Vol.22 No.1, pp.69-71. [9] Zhang Y.F., Mao J. L., An improved K-means Algorithm, Computer Application, vol.23. no.8, pp. 31-33,2019.