Microsoft Word - 4 Jiang Jianhua, Sheng Buyun, Yang Mingzhong--Research on OWA based Multi-Source Heterogeneous Data Fusion.do 232-239 Advances in Systems Science and Applications (2011), Vol.11, No.3-4 ISSN 1078-6236 International Institute for General Systems Studies, Inc. Research on OWA Based Multi-Source Heterogeneous Data Fusion Jianhua Jiang, Buyun Sheng and Mingzhong Yang School of mechanical and electronic engineering, Wuhan University of Technology, Wuhan 430070, China Abstract For the problem of multi-source heterogeneous data fusion, an architecture model of multi-source heterogeneous data fusion was designed. In the process of data fusion, triangular fuzzy number (TFN) was used to give uniform express of multi-data description in quantity, the ordered weight average (OWA) was used to deal with the preference of decision-maker and design the algorithm of data fusion. At last, the feasibility was validated by an example. Keywords data fusion; triangular fuzzy number; ordered weight average 1. Introduction Data fusion is a process of multi-source data cooperating as to reduce data redundancy and capture comprehensive information from data. It has become a research focus in data processing, object identification, situation assessment and intelligent decision-making domains. In general view, the data for data fusion come from multi-sensors[1], the data type is almost numerical and the methods are mainly from statistics and artificial intelligent, and some literatures have researched the multi-source data fusion[2,3,4]. In fact, except for numerical value, there are many other expression ways of data such as language and symbols, and the multi-ways of data expression will led to the ambiguity, difference and heterogeneity of data. On the other hand, in order to make an ultimate decision, a decision-maker must integrate a wide range of heterogeneous data and information. Therefore, in allusion to the features of heterogeneous data, this paper set focus on the method of multi-source heterogeneous data fusion and its application. 2. Multi-source Heterogeneous Data Fusion Methods and Its Architecture 2.1 multi-source heterogeneous data fusion methods The methods for data fusion in decision level mainly include weight average[5], D-S evidence theory[6] and voting[7]. (1) Weight Average Method The method uses formula ∑ iji tw to compute decision support value, where iw denotes weight of data source i and ijt denotes support value of data source i to decision j . It evaluates every decision according the support value, it is easy to operate and gives full consideration to the importance of data sources. But it is difficult to eliminate the impact of subjective factors in determining the weights. (2) D-S Evidence Theory It defines the space including all possible results of the object to be identified as a framework set D , its subset marked D2 and for DA ⊆∀ : ]1,0[2: →Dm Where 0)( =φm , 1)(2 =∑ ⊆ DA Am , φ is empty set, and m is called basic probability assignment function (BPAF) which actually refers to assign the trust value to subsets of D on Advances in Systems Science and Applications (2011), Vol.11, No.3-4 233 ISSN 1078-6236 International Institute for General Systems Studies, Inc basis of the available evidence. In practice, many BPAFs should be combined for different evidences of a problem may led to different im , it can be realized by the following formal: )()( 1 i AA i AmKAm i ∑∏×= =∩ − ( ni ≤≤1 ) Where )( i A i AmK i ∑∏= ≠∩ φ and DAi ⊆ D-S evidence theory is based on BPAF and able to compute uncertainty caused by unknown factors. But it requests all the items in D mutually exclusive and is difficult to calculate when there are too many BPAFs. (3) Voting It treats every data source as a voter and chooses a decision by comparing the votes obtained, the votes is defined as follow: ))(()( iji aSupFaSup = Where ia denotes decision i ( ni ≤≤1 ), )( iaSup denotes the total votes it obtained, )( ij aSup denotes the support value of data source j to decision i and its value is 0 or 1, and F may be defined as sum. It is difficulty to find BPAF for D-S evidence theory and distinguish two decisions have same votes, this paper uses OWA to fusion multi-source data in full considering the preference of decision-makers. 2.2 Multi-source Heterogeneous Data Fusion Architecture Literature [2] proposes an architecture for multi-source data fusion as shown in figure 1, it takes into account user requirements and source credibility, and uses proximity knowledge base, knowledge of reasonableness and voting method to resolve data conflicts. Figure 1 Schematic of data fusion Under the guidance of the above model, an architecture model of multi-source heterogeneous data fusion in decision level was designed as shown in figure 2. The fusion engine is made up of four parts of data warehouse, decision support value computing, OWA operator weight vector computing and data conversion and sorting. Figure 2 multi-source data fusion architecture (1) Data warehouse is used to integrate multi-source data and eliminate data heterogeneity and differences by operation of data selection, feature extraction and statistical analysis. 234 Jiang:Research on OWA based Multi-Source Heterogeneous Data Fusion (2) Decision support computing module obtains related data from data warehouse according to decision-making attributes and computes the support value ijs of data source i to decision j . (3) OWA operator weight vector computing module computes the weight of data source according to the fuzzy semantic parameters provided by decision-maker, and these parameters reflect the preferences of decision-makers to the data source. (4) Data conversion and sorting calculates ' ijs in combining with data source credibility (or importance degree) and OWA weight, and sorts ' ijs according its value. At last, the final decision value can be calculated by ' ijs and iw . 3. Algorithm of Data Fusion 3.1. Data Type and Its Features Data can be expressed in quantity and quality, and numbers for quantity while language variables for quality[8]. According to difference of data description, data type can be divided into quantitative and qualitative two types, this paper focus on 4 types of data description, as shown in table 1. Table 1 data type type Data Description notes Random variable Random variable submits to certain distribution quantitative Two value type Data value either 1 or 0, or true or false Degree type The degree general described by 7 or 9 standard qualitative Terminology based Description method depends on vocabulary space In case of large samples, random variable submits to normal distribution marked ),(~ 2σμX , where μ denotes expectation and σ denotes standard variance, and 9974.0)33( =+<<− σμσμ XP . Two-value type data is used to give answer for affirming or denying, if agreed the value is 1, otherwise 0, it can also be described by true or false. Degree type data is generally described by degree adverbs such as good, a little good and so on, the grading standard widely used is 7 or 9. Terminology based data is described by terms from a pre-defined vocabulary space and the space depends on specific circumstance. 3.2 Computing of TFN based Support Value In allusion to the ambiguity in data description, TFN is used to compute the decision support value. (1) Transform Method for Random Data Supposed that: σ30 −= ux σ6 0' xxx − = If support value increases with the increased of random variable x and ]3,3[ σμσμ +− is divided into n intervals, support value can be defined as: ⎪ ⎪ ⎩ ⎪⎪ ⎨ ⎧ +> + + ≤<+ + −≤ = σ σσ σμ 3,)1,1,1( )1(66,)1,,( 3,)0,0,0( )( 00 ' ux x n ixx n i n ix n i x xs (1) Advances in Systems Science and Applications (2011), Vol.11, No.3-4 235 ISSN 1078-6236 International Institute for General Systems Studies, Inc If support value decreases with the increased of random variable x , support value defined as: )()1,1,1(' )( xss x −= (2) Transform Method for Two Value Type Data Supposed that two-value type data described by set {1, 0}, the number of answer 1 and 0 is n and m respectively. If support value is defined by number of value 1, support value can be defined as: )/,/,/()( mnnmnnmnnxs +++= (2) (3) Transform Method for Degree Type Data If data described by the 7 grading standard, the support value can be defined according to table 2. Table 2 Fuzzy quantification of degree type data (4) Transform Method for Terminology based Data If the vocabulary space w includes n terminologies, and they are ordered by their contribution value for decision from low to high as },...,,{ 110 −= nwwww , the support value defined as: ))1/(),1/(),1/(()( −−−= nininiws i (3) 3.3 Computing of OWA Operator[9] Weight Vector Supposed that RRF n →: and a n dimension vector ),...,,( 21 nwwww = related to F makes: ∑ == n i iin bwaaaF 121 ),...,,( (4) Where ]1,0[∈iw , ni ≤≤1 and 11 =∑ = n i iw . ib is the ith largest element of ia . Then F is called n dimension OWA Operator and ),...,,( 21 nwwww = can be represented by a function as: )/)1(()/( nifnifwi −−= (5) Where ni ,...,2,1= , f is named fuzzy semantic operator (FSO) and defined as: ⎪ ⎩ ⎪ ⎨ ⎧ < ≤≤−− < = xb bxaabax ax xf ,1 ,)/()( ,0 )( (6) Where ]1,0[,, ∈bax . In addition, OWA also defines a measure operator for reflecting preference of decision-makers named )(worness and it is defined as: Inverse proportion Direct proportion )(xs very high very low (0.00,0.00,0.17) high low (0.00,0.17,0.34) a little high a little low (0.17,0.34,0.50) common common (0.34,0.50,0.67) a little low a little high (0.50,0.67,0.84) low high (0.67,0.84,1.00) very low very high (0.84,1.00,1.00) 236 Jiang:Research on OWA based Multi-Source Heterogeneous Data Fusion ∑ = −== n i iwin n wornessc 1 )(1)( (7) 3.4 Algorithm of Data Fusion Supposing there are n decisions ),...,,( 21 nAAAA = , m data sources ),...,,( 21 mSSSS = and ip is the credibility or importance of iS , the algorithm of data fusion is described as: Step1: Compute support value of data sources to decisions Access data from data warehouse and transform it to support value of every decision according to the method of section 3.1. The support value marked as: ),,( ijijijij cbaS = Where ijS is support value of data source i to decision j , ),,( jiijij cba is TFN expression of ijS and 10 ≤≤≤≤ ijijij cba . Step2: Compute OWA operator weight vector Considering the preference of decision-makers, choose fuzzy semantic parameters for formula 6. In most cases, fuzzy semantic parameters is defined as “the majority”, “at least half” and “as much as possible”, their parameters are )8.0,3.0( , )5.0,0( and )1,5.0( respectively. According to the parameters, the FSO )(xf can be determined. With FSO, formula 5 and 7, OWA weight vector ),...,,( 21 nwwww = and )(worness can be calculated. Step3: Transform ijs according to ip and ijs In order to use OWA weight vector, ijs must be transformed and sorted, the transform method as follow: Supposed that: ijiij sps ×=min_ ijiijiij spsps ×−+=max_ (8) ijin i i averageij sp p ns ∑ = = 1 _ Define: If 5.0≤c 0)( =ch , cm c 2)( = , cl c 21)( −= If 5.0≥c 12)( −= ch c , cm c 22)( −= , 0)( =cl ijs can be transformed into ' ijs by following formula: min_)(_)(max_)( ' ijcaverageijcijcij slsmshs ++= (9) Step 4: Compute final support value for every decision The final support value of decisions can be computed by the following formula: ∑ == m i ijij bws 1 nj ,...,2,1= (10) Where ijb denotes the ith largesse element of ),...,,( '' 2 ' 1 njjj sss . Step5: Make a decision according to the final support value of decisions. 4. An Example This paper takes steam turbine product development decision-making as example. Supposing there are 5 kinds of product named 1A , 2A , 3A , 4A and 5A , the data can be collected include market demand, product feedback, product parameters, historical data, failure statistics, Advances in Systems Science and Applications (2011), Vol.11, No.3-4 237 ISSN 1078-6236 International Institute for General Systems Studies, Inc expert advice and so on. Based on these data sources, products can be compared from 6 aspects, they are assessment of market demand 1a , average failures per year 2a ( 8.0,5.3 == σμ ), longest average no failure time 3a (month as unit and 53.2,28.12 == σμ ), economy evaluation 4a , customer evaluation 5a and expert advice 6a , and the preliminary results are shown in table 3. Table 3 support value of product and data source credibility (1) In table 3, 1a and 4a are degree-type data and transformed by table 2, 2a and 3a are random data and transformed by formula 1 (supposing n=15), 5a is two-value type data(the value in table 3 is the ratio of positive evaluation) and transformed by formula 2, 6a is terminology based data and transformed by formula 3. The transformed results are shown in table 4. (2) If “the majority” is selected as FSO, the parameters a and b in formula 6 are 0.3 and 0.8 respectively. According to formula 5 and 6, the OWA weight vector can be calculated as: )0,27.0,33.0,33.0,067.0,0(=w Put every iw into formula 7, the value of c can be calculated and 37.0=c . (3) According to formula 8 and 9, the data in table in table 4 can be transformed as shown in table 5. (4) For every column data in table 5, sorted them from high to low according the second value and computed final support value by formula 10, the result is shown in table 6. (5) From table 6, product 3A gets the highest support value, so it should be developed full Table 4 disposed attribute values ia ip A1 A2 A3 A4 A5 1a 0.5 (0.67,0.84,1.00) (0.34,0.5,0.67) (0.84,1.00,1.00) (0.17,0.34,0.50) (0.00,0.17,0.34) 2a 0.9 (0.13,0.14,0.20) (0.00,0.00,0.00) (0.53,0.55,0.60) (0.80,0.85,0.86) (0.02,0.23,0.27) 3a 0.9 (0.20,0.24,0.26) (0.53,0.56,0.6) (0.66,0.68,0.73) (0.13,0.16,0.20) (0.80,0.83,0.86) 4a 0.3 (0.34,0.50,0.67) (0.84,1.00,1.00) (0.17,0.34,0.50) (0.50,0.67,0.84) (0.67,0.84,1.00) 5a 0.6 (0.65,0.65,0.65) (0.73,0.73,0.73) (0.48,0.48,0.48) (0.54,0.54,0.54) (0.67,0.67,0.67) 6a 0.4 (0.33,0.33,0.33) (1.00,1.00,1.00) (0.66,0.66,0.66) (0.00,0.00,0.00) (0.33,0.33,0.33) ia ip A1 A2 A3 A4 A5 1a 0.5 large common Very large A little small Very small 2a 0.9 5.25 8.64 3.28 1.81 4.76 3a 0.9 8.36 13.28 15.14 7.08 17.31 4a 0.3 commo n Very good A little bad A little good good 5a 0.6 0.65 0.73 0.48 0.54 0.67 6a 0.4 Improve Fully recommen d As usual Out of market Improve 238 Jiang:Research on OWA based Multi-Source Heterogeneous Data Fusion Table 5 results of data transformed ia A1 A2 A3 A4 A5 1a (0.50,0.62,0.75) (0.25,0.37,0.50) (0.63,0.75,0.75) (0.13,0.25,0.37) (0.00,0.13,0.25) 2a (0.18,0.19,0.27) (0.00,0.00,0.00) (0.71,0.74,0.81) (1.07,1.14,1.16) (0.27,0.31,0.36) 3a (0.27,0.32,0.35) (0.71,0.75,0.81) (0.89,0.92,0.98) (0.18,0.22,0.27) (1.07.1.12,1.16) 4a (0.15,0.22,0.30) (0.38,0.45,0.45) (0.07,0.15,0.22) (0.22,0.30,0.38) (0.30,0.38,0.45) 5a (0.58,0.58,0.58) (0.65,0.65,0.65) (0.43,0.43,0.43) (0.48,0.48,0.48) (0.60,0.60,0.60) 6a (0.20,0.20,0.20) (0.60,0.60,0.60) (0.39,0.39,0.39) (0.00,0.00,0.00) (0.20,0.20,0.20) Table 6 final support values A1 A2 A3 A4 A5 value (0.23,0.27,0.31) (0.43,0.49,0.53) (0.52,0.54,0.56) (0.20,0.27,0.35) (0.28,0.32,0.36) 5. Conclusion In this paper, an architecture model for multi-source heterogeneous data fusion was constructed, TFN based data processing for multi-data description in quantity was researched and OWA operator based data fusion algorithm was designed, the practical application indicates that the algorithm designed is effective. The study of this paper presents a feasible option for constructing intelligent decision support system and has certain reference value for similar data processing and fusion. Acknowledgements This paper is supported by National Natural Science Fund Project of China (contract no.50620130441), Scientific and Technological Project of Wuhanc City (contact no. 200810321153) and Youth Science and Technology Chenguang Project of Wuhan City (contact no. 200750731289). References [1] D. L. Hall, J. Llinas. Hand Book of Multi-sensor Data Fusion. Beijing: Electronic Industry Press, 2008. [2] R. R.Yager. A framework for multi-source data fusion. Journal of Information Sciences Vol.163 (2004) 175~200. [3] Zhou Haiyin, Li Donghui, Jiang Yueping. High Precision Method for Determining the Position of Aerocraft Based on Mult-Sensors Information Fusion. Journal of Hunan University(natural Sciences), 34(4) (2007) 37~40. [4] R.R.Yager. Modeling Intelligence information: Multi-source fusion. IEEE International Conference on Computational Intelligence for Homeland Security and Personal Safety,Orlando, FL, MAR 31-APR 01, 2005. [5] Wang Guangyun, Li Weihua, Hua Wenjian, etal. A method for heterogeneous uncertain information fusion and its application. International Conference on Signal Processing Proceedings, vol.3 (2004) 2253~2256. [6] Ma Linru, Yang Lin, Wang Jianxin. Research on Security Information Fusion from Multiple Heterogeneous Sensors. Journal of System Simulation, 20(4) (2008) 981~985. [7] S. Kang. Two-Phase Identification Algorithm Based on Fuzzy Set and Voting for Intelligent Multi-sensor Data Fusion, Lecture Notes in Computer Science, vol. 4252 (2006) 769~776. [8] F.Herrera, L.Martinez. An approach for combing linguistic and numerical information based on the 2-tuple fuzzy linguistic representation model in decision-making. International Advances in Systems Science and Applications (2011), Vol.11, No.3-4 239 ISSN 1078-6236 International Institute for General Systems Studies, Inc Journal of Uncertainty, Fuzziness and knowledge-based systems, 8(5) (2000)539~562. [9] Z.Zhang, C.Zhang. Result fusion in multi-expert systems based on OWA operator. Proceedings of the 23rd Computer Science Conference, (2000) 234~240.